Content
By default, Docker containers run as the root user. This means if an attacker exploits a vulnerability in your application, they gain root access inside the container. Combined with a container escape vulnerability, this could translate to root access on the host machine . The principle of least privilege demands that containers run with only the permissions necessary to perform their function.
The fix is straightforward: create a non-root user in your Dockerfile and switch to it using the USER instruction. Many official base images already include a non-root user—for example, Node.js images include a node user with UID 1000 . If your base image does not provide one, create it yourself:
dockerfile
RUN addgroup -S appgroup && adduser -S appuser -G appgroup
RUN chown -R appuser:appgroup /app
USER appuser
For an additional layer of defense, Docker's rootless mode runs the Docker daemon and containers as an unprivileged user, mitigating vulnerabilities in the daemon itself .
Mistake 2: Using Vulnerable or Bloated Base Images
Every base image brings its own set of installed packages, each with its own vulnerability history and patch cadence. A full Ubuntu image ships with shells, package managers, and utilities that an application container rarely needs at runtime—but each is a potential attack vector . Log4Shell demonstrated how far a single vulnerable dependency baked into base images can spread .
Minimal base images such as Alpine, Distroless, and Wolfi reduce the attack surface by stripping out everything except what is needed to run your application . Multi-stage builds reinforce this by keeping compilers and build tools entirely out of the final image.
Pinning base images by digest rather than tag is equally important. A mutable tag like python:3.11 can point to a different image tomorrow, silently altering what you ship . Use FROM python:3.11@sha256:<digest> to guarantee reproducibility.
Mistake 3: Exposing Unnecessary Ports
Docker publishes container ports using iptables destination NAT on the FORWARD path, which means host-side firewall rules on the INPUT chain do not catch them . A container with ports: - "5432:5432" for a database is reachable from any network that can route to the host, regardless of what the host firewall says.
Omit ports entirely when a service only needs to communicate with other containers on the same Docker network . Use expose instead, which makes the port available to linked containers without publishing it to the host. If a service must be reachable externally, bind to a specific interface rather than 0.0.0.0, and use internal: true on networks that need no external connectivity.
Mistake 4: Leaking Secrets into Image Layers
A secret written into a Docker layer is recoverable forever, even if a later layer deletes the file. The delete happens in one layer; the secret remains in the earlier one, extractable with docker history or by unpacking the image tarball . Build arguments and environment variables are equally dangerous: both are recorded in image metadata and visible to anyone with access to the image .
The correct approach for build-time secrets is BuildKit's secret mount, which exposes a secret to a single RUN instruction without writing it to any layer:
dockerfile
RUN --mount=type=secret,id=npmrc,target=/root/.npmrc \
npm ci --omit=dev
Runtime secrets should be injected at container start via Docker secrets, mounted files, or an external secret manager—never embedded in the image .
Mistake 5: Granting Excessive Container Privileges
Docker containers start with a default set of Linux capabilities that most workloads never use. The --privileged flag disables container isolation almost entirely, including capability restrictions, device cgroups, and seccomp . Mounting the Docker socket into a container is equally dangerous: it hands that container full control of the Docker daemon, which runs as root on the host .
Drop all capabilities and add back only what the process needs:
bash
docker run --cap-drop=ALL --cap-add=NET_BIND_SERVICE \
--security-opt=no-new-privileges myapp:secure
The no-new-privileges flag prevents setuid binaries inside the container from ever gaining more privilege than the process started with . Never mount /var/run/docker.sock into an untrusted container; use a rootless builder like BuildKit or Kaniko if a workload needs to build images .
Mistake 6: Skipping Image Vulnerability Scanning
An image is not just your application code—it is an entire filesystem with OS packages, language runtimes, and dependencies, each with its own vulnerability surface . A base image that scanned clean three months ago may now carry critical CVEs because vulnerability databases change daily .
Integrate scanning into your CI pipeline using tools like Trivy, Grype, or Docker Scout. Scan the built image rather than the Dockerfile, since only the assembled image reflects the full set of installed packages . Configure scans to fail the build on critical and high-severity findings, and run scans continuously as new CVEs are published .
Mistake 7: Using a Writable Root Filesystem
A writable root filesystem gives an attacker who compromises your application the ability to drop a webshell, modify binaries, or install persistence mechanisms . A read-only root filesystem stops these post-exploitation activities cold.
Run containers with --read-only and mount specific writable paths as tmpfs or volumes:
bash
docker run --read-only \
--tmpfs /tmp:rw,noexec,nosuid,size=64m \
myapp:secure
The noexec and nosuid flags on the tmpfs mount prevent execution of dropped binaries and disable setuid behavior on that filesystem . Apply patches by rebuilding and redeploying images rather than modifying running containers, which are ephemeral by design .
Mistake 8: Omitting Resource Limits
A container with no resource limits can exhaust host memory, CPU, or process tables and take down neighboring containers—a denial-of-service risk that is entirely preventable . Set explicit bounds for memory, CPU, and process count:
bash
docker run --memory=512m --memory-swap=512m \
--cpus=1.0 --pids-limit=200 myapp:secure
These limits prevent a single compromised or misbehaving container from affecting the availability of others on the same host.
Before-and-After Dockerfile
Before (insecure):
dockerfile
FROM node:20
WORKDIR /app
COPY . .
ARG NPM_TOKEN
RUN echo "//registry.npmjs.org/:_authToken=$NPM_TOKEN" > .npmrc && \
npm ci && rm .npmrc
ENV DATABASE_URL=postgres://admin:password@db:5432/app
EXPOSE 3000
CMD ["node", "src/index.js"]
This Dockerfile runs as root, uses a full base image, bakes the npm token and database credentials into layers, and exposes the application port broadly.
After (hardened):
dockerfile
# syntax=docker/dockerfile:1
FROM node:20-bookworm-slim@sha256:abc123... AS builder
WORKDIR /app
COPY package*.json ./
RUN --mount=type=secret,id=npmrc,target=/root/.npmrc \
npm ci --omit=dev
COPY . .
RUN npm run build
FROM node:20-bookworm-slim@sha256:abc123...
WORKDIR /app
COPY --from=builder /app/node_modules ./node_modules
COPY --from=builder /app/dist ./dist
COPY --from=builder /app/package.json ./
RUN addgroup -S appgroup && adduser -S appuser -G appgroup && \
chown -R appuser:appgroup /app
USER appuser
EXPOSE 3000
CMD ["node", "dist/index.js"]
The hardened version uses a pinned slim base image, a multi-stage build to keep build tooling out of the final image, BuildKit secret mounts for the npm token, a non-root user, and no embedded runtime secrets. Database credentials are injected at runtime via Docker secrets or environment variables from a secure source.