Documents the current pipeline anatomy, the metrics to capture on the first real run, hypothesised bottlenecks, and ranked optimisation proposals. Placeholders for real numbers; to be filled in after the first deploy lands. Companion to the Portainer redeploy step in .gitea/workflows/ci-cd.yaml on the same branch.
14 KiB
CI/CD optimization proposal — davidaragon-portfolio
Status: Draft v0.1 — initial scaffold, real measurements pending.
Last updated: 2026-07-10
Owner: David Aragón + Hermes
Related: docs/architecture.md · docs/ci-cd-setup.md
0. Why this document exists
The CI for davidaragon-portfolio is the second pipeline brought into the
Gitea Actions + Portainer setup on the NAS (the first being
python-project-template-internal). Both share the same QNAP-hosted
act_runner self-hosted runner and the same Portainer instance — that means
an optimization here is partly informed by what we already learned on the
Python template, and partly needs its own measurements because the build is
different (Node 20 + Astro static vs Python 3.10 + uv).
The first real deploy of the portfolio stack is the gating event: until that
runs and we can read the timing breakdown, this document is structural +
conjectural for any section that needs real numbers. Once we have one
truthful run, the open [METRIC] placeholders get filled in.
1. Anatomy of the current pipeline
The workflow lives at .gitea/workflows/ci-cd.yaml and currently has:
push / pull_request / workflow_dispatch(deploy:boolean)
|
+--> job: build (runs-on: ubuntu-latest)
|
+--> step 1: actions/checkout@v4
+--> step 2: docker login + build + push
| (tags: :latest and :0.0.1)
+--> step 3 (conditional): Redeploy stack on Portainer
|
+--> a) DELETE existing stack on Portainer API
+--> b) POST a fresh stack from docker-compose.prod.yml
+--> c) sleep 8s, then smoke-test http://host:3001/
The runner
runs-on: ubuntu-latest— noself-hostedlabel required.- The QNAP-resident
gitea/act_runner:latestregisters with these labels (per the existing CI config that runs successfully today):ubuntu-latestlinuxx64- probably
act-runner(informally used internally)
- Single shared runner. Iteration
iter 5of the Python template proved that withoutconcurrency:group withcancel-in-progress: false, two parallel runs cancel each other's mid-flight steps. This workflow inherited that fix from PR #2 (seef152379 fix(ci): add concurrency guard + PR trigger + workflow_dispatch). No regression expected on this repo.
What runs on each step
| Step | Container / binary used | On the runner | Notes |
|---|---|---|---|
| 1 (checkout) | actions/checkout@v4 (Node) |
Yes | Adds ~1s |
| 2 (build) | local docker buildx |
Yes | Pulls node:20-alpine + nginx:alpine. Network bound. |
| 3 (deploy) | local curl + jq |
Yes | Calls Portainer API at http://${PORTAINER_URL}/api/stacks/.... Internal network bound. |
The runner must therefore have:
- Outbound HTTPS to
registry.npmjs.org(during build, npm ci) - Outbound HTTPS to
gitlab.impresion3d.pro:443(during push) - Network reachability to the Portainer URL (
PORTAINER_URLsecret). When that URL ishttp://192.168.1.30:9000, the runner needs to be on the same LAN as the NAS. The QNAP self-hosted runner satisfies this only if the runner container has access to the LAN (either via--network hoston the runner or via a docker network with a route to192.168.1.0/24).
The third point is the most likely silent failure: a --network=some-bridge
runner cannot resolve 192.168.1.30 by default and the redeploy step will
then fail with connection refused or No route to host. Before the first
real deploy, the runner's network scope should be verified — see the
verification checklist.
2. Critical metrics to capture on the first deploy
The first real run will fill these in. Once captured, copy them into this section and update the subsequent analysis.
| Metric | Source | Target | Actual (first run) |
|---|---|---|---|
t_checkout |
logs of step 1 | < 5 s | [METRIC] |
t_docker_pull_alpine_node |
logs of step 2 (first lines) | < 30 s | [METRIC] |
t_npm_ci |
step 2 lines mentioning npm ci |
< 60 s | [METRIC] |
t_astro_build |
step 2 lines mentioning astro build |
< 30 s | [METRIC] |
t_docker_build_total |
step 2 end timestamp - step 2 start | < 180 s | [METRIC] |
t_docker_push_latest |
step 2 lines mentioning latest push |
< 20 s | [METRIC] |
t_docker_push_pinned |
step 2 lines mentioning :0.0.1 push |
< 20 s | [METRIC] |
t_step_3_delete |
step 3 DELETE line | < 5 s | [METRIC] |
t_step_3_create |
step 3 POST line | < 5 s | [METRIC] |
t_step_3_smoke_total |
step 3 sleep + curl | ~8 s + connect time | [METRIC] |
| Total wall-clock (with deploy) | run-level duration | < 5 min | [METRIC] |
| Total wall-clock (without deploy) | run-level duration | < 3 min | [METRIC] |
| Runner queue waiting time | run started - run created | < 60 s | [METRIC] |
To capture these, the easiest path is to download the job logs via the Gitea API once the run is complete:
# Get run_id from the Gitea UI (or /api/v1/repos/.../actions/runs)
curl -s -H "Authorization: token $GITEA_TOKEN" \
"$GITEA_URL/api/v1/repos/root/davidaragon-portfolio/actions/runs/<RUN_ID>/jobs" \
| jq '.[].id'
# Then fetch logs per job
curl -s -H "Authorization: token $GITEA_TOKEN" \
"$GITEA_URL/api/v1/repos/root/davidaragon-portfolio/actions/jobs/<JOB_ID>/logs"
The gitea-actions-ci-debugging skill in the agent's repo has more on this
pattern (see references/api-endpoints.md).
3. Hypothesised bottlenecks (to be verified against §2)
The following are guesses to be confirmed or refuted with the first run. They are listed in order of expected impact.
H1. Docker layer cache cold on every run
- Why: The runner does not persist
docker buildxcache between runs unless a layer cache is configured. - Symptom:
t_docker_pull_alpine_nodelikely to be ~30s on cold cache. - Possible fix (sized for the runner's disk):
- Mount a host directory to
~/.cache/buildxand exportBUILDX_CACHEenv. - Or use
docker/build-push-action@v6withcache-from: type=registry.
- Mount a host directory to
- Effort / risk: low. Touches only step 2.
H2. Sequential :latest then :0.0.1 pushes
- Why: Image bytes are uploaded twice (once per tag), doubling push time.
- Fix: tag both and push only
:0.0.1, then:latestis left for ad-hoc inspection. Or use--all-tags(docker push --all-tags) which sends each unique manifest only once if the registry supports it (Gitea container registry supports it since 1.20). - Saving: ~50% of push time (typically < 10s on LAN).
H3. npm ci always re-downloads when package-lock.json changes
- Why: Without an
npmcache, every push with a new lock entry is a fresh download fromregistry.npmjs.org. - Symptom:
t_npm_ciproportional to dep delta. With the current lock file (Astro 5 + Tailwind + Zod), cold downloads are ~30s for the small tree, but as more plugins are added this grows. - Fix: use
actions/setup-node@v4withcache: 'npm'(it caches based onpackage-lock.jsonhash automatically). Or usecache@v4directly with pathnode_modulesand key derived from the lock hash. - Saving: 30s+ per run on a miss, ~5s on a hit.
H4. Image pull-from-registry at container runtime
- Symptom in step 3: the first health check at
:3001/may be slow on cold cache because the new container has to pull...:0.0.1from the registry. The compose file does not pinpull_policy: always, so the runner may use an older layer cache on the runner host. - Fix: add
pull_policy: alwaystodocker-compose.prod.ymlso the new image is fetched every time.
H5. Concurrency guard might be too coarse
- Why: The current group is
ci-${{ github.ref }}— per ref. This serialises pushes tomainAND pushes to PR branches on top of each other. With a single runner, this is correct. If a second runner comes online later, the guard may need to be re-evaluated. - No action for now. Re-evaluate when a second runner is registered.
4. Optimisation proposals (ranked by ROI)
These are proposals to apply after the first run, ranked by expected return on effort.
Proposal A — Add actions/setup-node@v4 with npm cache (effort: 5 min, expected saving: 30s/run on miss, 5s on hit)
Add as a new step between checkout and build:
- uses: actions/setup-node@v4
with:
node-version: '20' # matches Dockerfile
cache: 'npm' # caches ~/.npm based on package-lock.json hash
This costs nothing in code maintenance and helps every PR build.
Proposal B — Switch push to --all-tags (effort: 2 min, saving: ~10s/run)
Replace:
docker push gitlab.impresion3d.pro/root/davidaragon-portfolio:latest
docker push gitlab.impresion3d.pro/root/davidaragon-portfolio:0.0.1
with:
docker push --all-tags gitlab.impresion3d.pro/root/davidaragon-portfolio
Both tags get pushed in one HTTP exchange; the registry deduplicates identical blobs.
Proposal C — Add pull_policy: always to compose (effort: 1 min, saving: variable)
Add the line to docker-compose.prod.yml:
services:
portfolio:
image: gitlab.impresion3d.pro/root/davidaragon-portfolio:0.0.1
pull_policy: always # <-- new
restart: unless-stopped
ports:
- "3001:80"
Without this, the runner host may serve a cached older layer to the new container even though the tag on the registry has changed.
Proposal D — Drop the :latest tag entirely (effort: 5 min, saving: 10s/run, cost: 1 fewer tag for ad-hoc inspection)
Pure-only-:0.0.1 means every pull is explicit (CI passes the pinned tag).
Pros: zero ambiguity in docker images listings. Cons: less convenient for
docker exec -it foo /bin/sh quick explorations.
Not recommended unless the user explicitly asks — the :latest is useful as
a "what's currently live" pointer.
Proposal E — Squash the image with docker-slim or dive optimisation (effort: 1 hour+, saving: 20-40 MB image size)
Probably not worth it for a portfolio site. Skip.
Proposal F — Parallel runs (long-term)
Currently the pipeline has a single build job. If lint, type-check, and
build grow apart in cost, splitting them into parallel jobs would help. But
the Astro site has no type-checking step (vanilla JS), and a project this
small does not benefit from parallelisation. Leave as-is.
5. Open questions for the user
- Should the build always push both
:latestand:0.0.1, or only the pinned version? (Affects Proposal D) - Is there any reason to keep the smoke-test in step 3 (
http://host:3001/)? The actual public URL is via NPM (https://davidaragon.impresion3d.pro). Smoke-testing the public URL would be more accurate but requires the runner to have outbound internet AND DNS for the FQDN.
6. Risks of the current design
- Single runner bottleneck. The QNAP has exactly one
act_runnercontainer; pushes to two repos (template + portfolio) serialise. If both repos ever have long-running builds, developer iteration slows. No fix for now (registering a second runner is on the long-run roadmap). pull_policynot set. The compose does not force a pull-on-create. See Proposal C above.- No
if: failure()retry or notify step. If the Portainer redeploy fails, the run ends withfailurebut no notification pings anyone. The user has to look at the Actions tab manually. Add a Slack/email notification step in a future iteration (low priority while there is one user). - The
IMAGE_NAMEenv var in the workflow is hardcoded. When the version bumps from0.0.1to0.0.2, a separate edit is needed inci-cd.yaml. Could be moved to aVERSIONrepo variable later.
7. Verification checklist before first real deploy
Run through this list before firing workflow_dispatch(deploy=true) to
reduce surprise:
act_runnerisIdle(orOnline) inhttps://gitlab.impresion3d.pro/-/admin/actions/runners. Not inRestarting.act_runnerhas the labelsubuntu-latest,linux,x64at minimum.act_runner's docker network can resolve192.168.1.30(verify with a debug step that runsgetent hosts gitlab.impresion3d.pro).- Portainer URL responds from the runner (verify with a debug step:
curl -I ${PORTAINER_URL}/api/status). - All 4 secrets are visible in
https://gitlab.impresion3d.pro/root/davidaragon-portfolio/settings/secrets/actionswith values(hidden). DOCKER_USERNAMEandDOCKER_PASSWORDwork (verify withdocker login ... -u $U -p $P && docker push <empty-image>once).- The
act_runnerhost has at least 2-3 GB free disk (for buildx cache + pulled base images).
The first three items live at the NAS / runner level — see the
qnap-nas skill for the diagnostics. Items 4-7 are repo-level.
8. Proposed workflow to apply A, B, C
After the first deploy succeeds and §2 is filled in, apply A, B, C in a
single PR. The PR title can be ci: apply optimisations A B C from docs/ci-optimization-proposal.md. Estimated effort: 10 minutes. Estimated
saving: ~30-50 s per run.
If §2 reveals a metric that disagrees with §3 (e.g. push is actually fast
and astro build is the slow part), the proposal list should be re-ranked
before applying, not after.
Appendix A — Convention: when to write a separate skill for this
If the same discovery shows up 3 times across this repo and the Python template, it graduates from "proposal" to "skill". Candidates so far:
- Image tag strategy for self-hosted GitLab Container Registry — both repos push to the same registry. Worth a skill once the second project ships the same pattern.
- Concise Portainer DELETE+CREATE pattern — the bash snippet in step 3 is reusable. Worth a skill the second time it's copy-pasted.