The pipeline, under the hood

In the last post the pipeline was a footnote: push to main, and some forty seconds later the post is live. This one is about what happens during those seconds — and, more to the point, why it is built this way rather than some other way.

The goal was narrow from the start: a git push should update the site without me logging into the server. No SSH, no hand-run docker compose, no folder to upload.

Why an image and not a file sync

The obvious solution would have been to copy the built dist/ folder to the server with rsync or scp. Fewer moving parts, faster to set up. I decided against it, for three reasons.

A file sync is not atomic. While rsync runs, old and new files sit side by side on the server. With a static site using hashed asset names this is usually harmless — but only usually, and an HTML document pointing at a stylesheet that hasn’t arrived yet is exactly the kind of failure you can never reproduce afterwards.

A file sync has no memory. Once the copy is done, the old files are gone. Rolling back means checking out the previous state locally, rebuilding, uploading again. An image tagged with a commit SHA, by contrast, is a named, immutable artefact. Going back to the state from two days ago means changing a tag and recreating the container.

And it doesn’t match the rest. The applications planned for later — a Python app, a Spring Boot app — need images regardless. Maintaining two different deployment mechanisms on the same server would have been the worse trade, made purely to save a few lines on the simplest application.

The Dockerfile: two stages

# Stage 1: build the site
FROM node:22-alpine AS build
WORKDIR /app
COPY package.json package-lock.json ./
RUN npm ci
COPY . .
RUN npm run build

# Stage 2: serve the finished files only.
FROM nginx:alpine
COPY --from=build /app/dist /usr/share/nginx/html

Two details matter more than they look.

npm ci instead of npm install, with the lockfile copied explicitly ahead of the rest of the source. Together: the build pulls exactly the versions from the lockfile and nothing else, and the dependency layer is only rebuilt when the lockfile actually changed. npm install is allowed to modify the lockfile when it thinks it has found a better resolution. In a CI run that isn’t convenience, it’s a silent divergence between what was tested locally and what gets shipped.

The second detail sits in the Dockerfile as a comment block, because otherwise someone — most likely me, six months from now — will mistake it for a bug: the image deliberately ships no nginx configuration of its own.

The target container on the server has its conf.d directory bind-mounted from the host, along with the certificates and the redirect rules. A Docker bind mount replaces the target directory in the image completely — had I baked a config into the image, it would have been silently shadowed at startup. No error, no warning, just permanent confusion about why changes have no effect. Port, TLS and the listen and root directives remain the server’s business. The image ships content.

The elegant side effect: nginx:alpine comes with its own default.conf, which happens to point at exactly /usr/share/nginx/html. Locally, that makes the image startable and testable on its own, as if it were a complete web server. On the server, that same file is covered by the host mount. One image, two roles, with no special case in the build.

A custom validation script hangs off package.json as a prebuild hook. npm runs it automatically ahead of every npm run build — among other things, it checks that every post exists in both languages under the same filename.

That it is a lifecycle hook rather than a separate workflow step is the point. The Dockerfile’s first stage calls npm run build, so the check runs inside the image build. If a translation is missing, the Docker build fails and no image is produced at all — regardless of whether the pipeline is building or someone is building locally, bypassing it. A rule that lives in an author’s head is not a rule; one you can only bypass by editing the Dockerfile is.

Job 1: build

permissions:
  contents: read
  packages: write

These three lines sit at the top of the workflow file and are the part most easily left out. Without them the automatically provided GITHUB_TOKEN inherits the repository’s default permissions — considerably more than a build needs. With them, the token may read source and write packages. Nothing else.

That it is the GITHUB_TOKEN at all, rather than a personal access token, is the second half of the same decision. The token is minted per run, expires with the run, and cannot be accidentally reused elsewhere. A PAT in the secret store would be a long-lived credential that someone would eventually have to rotate — and that nobody then rotates.

- name: Lowercase the image name
  run: echo "IMAGE=ghcr.io/${GITHUB_REPOSITORY_OWNER,,}/odabas-portfolio" >> "$GITHUB_ENV"

This step looks like housekeeping and is in fact a bug fix. The GitHub Container Registry only accepts lowercase image names. My GitHub account is WaveRider52, with two capitals. Using ghcr.io/${{ github.repository_owner }}/odabas-portfolio directly — the route almost every tutorial shows — therefore fails on push, with an error message that doesn’t exactly spell out the reason.

The fix is Bash parameter expansion: ${VAR,,} lowercases the contents. It works because the steps run in Bash on a Linux runner — GitHub Actions’ own ${{ }} expression syntax has no equivalent. The value is then handed to subsequent steps via $GITHUB_ENV, instead of being written out three times.

The rest of the job is deliberately unremarkable: the official actions for checkout, registry login and build-and-push. Tagging happens twice, with latest and with the full commit SHA.

tags: |
  ${{ env.IMAGE }}:latest
  ${{ env.IMAGE }}:${{ github.sha }}

The duplication is intentional. latest is what the Compose file on the server references, which means the deploy step never has to touch that file. The SHA tag is what makes a rollback possible at all. Pushing only latest would be a deployment without history: you always know what is running right now, but no longer what was running before.

Job 2: deploy

The second job depends on the first via needs and starts by setting up SSH access:

- name: Prepare SSH
  env:
    SSH_KEY: ${{ secrets.VPS_SSH_KEY }}
    KNOWN_HOSTS: ${{ secrets.VPS_KNOWN_HOSTS }}
  run: |
    install -m 700 -d ~/.ssh
    printf '%s\n' "$SSH_KEY" > ~/.ssh/id_ed25519
    chmod 600 ~/.ssh/id_ed25519
    printf '%s\n' "$KNOWN_HOSTS" > ~/.ssh/known_hosts

Three small things in there are deliberate.

install -m 700 -d creates the directory with the right permissions immediately. The alternative — mkdir followed by chmod — leaves a directory with overly open permissions for a moment. On a throwaway runner that is inconsequential, but it is the kind of habit worth forming correctly the first time.

printf '%s\n' rather than echo: depending on the shell, echo interprets backslash sequences. An SSH key is multi-line base64 text, and a tool that second-guesses it is the last thing you want here. printf '%s' emits the string unchanged, full stop.

And the secrets travel through the env: block rather than directly via ${{ }} inside the script line. Anything inside ${{ }} is substituted into the script text before execution — the key material would become part of the shell script itself. Through env: it lands in an environment variable and is never interpreted as code.

Then comes exactly one command:

ssh "$TARGET" "cd ~/projects && docker compose pull nginx-odabas && docker compose up -d --force-recreate nginx-odabas"

The --force-recreate is not optional here, however redundant it looks. docker compose up -d compares the service definition against the running container. The definition is unchanged — same name, same :latest tag reference — so Compose sees no reason to act. The tag stayed the same; the image behind it did not. Without --force-recreate the old container keeps running happily, the job reports success, and the site does not change.

What the command deliberately omits: any nginx -s reload, any touching of the reverse proxy, any handling of certificates. The edge router picks up the new container from its labels and keeps routing. The container name stays the same, so the certificate renewal hooks attached to that name stay untouched as well. The swap itself takes under a second.

A variable, not a secret

if: ${{ vars.DEPLOY_ENABLED == 'true' }}

DEPLOY_ENABLED is a repository variable, not a secret. This isn’t cosmetic. A secret is masked in the logs — a masked switch makes debugging harder while protecting nothing, because “deployment is turned on” is not a secret. Secrets are for things where knowing them grants access. Everything else belongs in variables.

The practical benefit was larger than expected. The first runs happened with the deploy step switched off: the image landed in the registry, the tagging could be verified, the build job went green — and the server stayed completely untouched. Those runs are still visible in the run history with the deploy job skipped. Only once everything else was verified did the switch get flipped.

A deploy step you arm after rehearsing it without risk is a considerably calmer thing than one whose first run is also its first live fire.

The deploy key

The pipeline has its own SSH key, not the one I log in with. It was generated directly on the server; the public half lives there in authorized_keys, the private half sits as a repository secret.

The reason is containment. A CI key by its nature lives in a system I do not fully control. If it is ever compromised, I want to be able to revoke that specific key — not lose my own access alongside it and then work out how to reach a server without access in order to repair access.

A side effect of generating it on the server: the private key never sat on my laptop and never travelled through an intermediate step.

The secret people leave out

The deploy job needs four secrets: host, user, private key — and the known host keys from ssh-keyscan.

The fourth is the one tutorials like to skip. Without it the pipeline has to disable host key verification, usually with StrictHostKeyChecking=no. That is the check which prevents something other than your own server from presenting itself as the destination. An automated deploy that accepts any host is an automated deploy to some host — carrying a private key with it.

Storing the output of ssh-keyscan -H costs a minute. Disabling the check saves that minute and trades it for a weakness nobody notices, because nothing breaks.

Even a working pipeline ages

The runs are green, and there are still three annotations underneath them. They are the most interesting part of the build log.

The first: actions/checkout@v4, docker/login-action@v3 and docker/build-push-action@v6 all declare Node 20 as their runtime — a version GitHub has now deprecated. The runner therefore already executes them on Node 24. That is more notable than it sounds: those actions are running today on a runtime their authors never tested them against, because the one they declare no longer exists. It works, but it works on grace.

The second and third apply to both jobs: the ubuntu-latest label moves to Ubuntu 26 starting 19 October 2026.

Neither breaks anything today, and both are the same mechanism I already rejected for the image: a moving reference. ubuntu-latest is the runners’ :latest — same name, eventually a different operating system. And @v4 is not a fixed version but a tag that travels along within its major line; which Node runtime sits behind it is decided not by my workflow file but by wherever that tag currently points.

The difference between a stable Friday afternoon and an inexplicably broken one is often just that something changed behind an unchanged name — and that you didn’t know when.

On the server side I have already drawn the conclusion: the Compose definitions are getting explicit version tags instead of latest. I have since done the same for the pipeline — runs-on: ubuntu-24.04 instead of the moving label, and the three actions bumped to their Node 24 majors. The next run showed no annotations at all. A pre-announced break you pin down in advance is maintenance; the same break three weeks later is firefighting.

What the pipeline deliberately does not do

The honest section, because a pipeline with no stated limits merely looks complete:

No automatic rollback. If the deploy fails or the site serves nonsense, nothing happens on its own. The SHA tag makes a manual rollback possible — that is all it is.

No post-deploy health check. The job counts as successful once the ssh command returns without error. “The container is running” and “the site works” are not the same statement, and the pipeline only knows the first one.

No staging environment. Builds go straight at production. For a static site with no database and no user data that is proportionate; for a stateful application it would be negligent.

No cleanup on the server. Old images accumulate until someone removes them.

These are known gaps, not overlooked ones — and that distinction is the actual point. The failure mode of this site is: the blog is broken for a few minutes and I revert a tag. Building automatic rollback with a health-check gate would be effort spent against a risk that does not exist. For the planned database-backed applications the arithmetic is different — and that is when it gets built, not beforehand on principle.

Result

The workflow responds to two triggers: every push to main and, via workflow_dispatch, the button in the Actions interface. The first run with the deploy step armed was triggered by hand — 49 seconds, both jobs green, the site confirmed in the browser afterwards. The push-driven runs before and after land between 37 and 52 seconds.

Since then, ordinary content updates work like this: write, commit, push. The server is no longer part of the workflow.

How the switch from the old installation to this architecture actually went — including the moment a single line in the Compose file was changed — is the subject of the next post.