{"article":{"slug":"inside-forge-writing-a-container-runtime-from-scratch","title":"Inside Forge: Writing a Container Runtime From Scratch","subtitle":null,"summary":"A from-scratch walkthrough of Forge, a container runtime built in Go on Linux primitives—covering namespaces, cgroups, rootfs setup, and lessons from building the runtime by hand.","content_type":"tutorial","language":"en","canonical_url":"https://stevenstank.vercel.app/blog/inside-forge-container-runtime-from-scratch","author":{"name":"saksham","url":"https://stevenstank.vercel.app","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"saksham","url":"https://stevenstank.vercel.app","listing_slug":null,"listing":null},"topics":[{"name":"Systems Programming","slug":"systems-programming","url":"https://listedarticles.com/topics/systems-programming"},{"name":"Open Source","slug":"open-source","url":"https://listedarticles.com/topics/open-source"},{"name":"Programming","slug":"programming","url":"https://listedarticles.com/topics/programming"},{"name":"Tutorials","slug":"tutorials","url":"https://listedarticles.com/topics/tutorials"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":3673,"reading_minutes":16,"published_at":"2026-08-30T12:00:00.000Z","added_at":"2026-09-29T18:12:18.596Z","updated_at":"2026-09-29T18:12:18.596Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":false},"profile_url":"https://listedarticles.com/articles/inside-forge-writing-a-container-runtime-from-scratch","markdown_url":"https://listedarticles.com/articles/inside-forge-writing-a-container-runtime-from-scratch.md","example":false,"citation":"saksham, saksham. \"Inside Forge: Writing a Container Runtime From Scratch.\" 30 Aug 2026. https://stevenstank.vercel.app/blog/inside-forge-container-runtime-from-scratch (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://stevenstank.vercel.app/blog/inside-forge-container-runtime-from-scratch"},"body_markdown":"# Inside Forge | Writing a Container Runtime From Scratch\n\n# i built my own container runtime in go\n\ni've used containers before. you run something, it works, and you don't really think about what is happening underneath. you just know that somehow this thing called a \"container\" makes your application run in its own environment.\n\nand honestly, i didn't really like that. i wanted to know what is actually happening when you run a container. what makes it isolated? how does it get its own filesystem? how does it get its own network? how does it stop a process from using all the memory on the machine?\n\nso instead of just reading about it, i decided to build one myself.\n\nthat's how **forge** started.\n\n## so, what is forge?\n\nforge is a small container runtime that i built from scratch in go. the source is on github ↗\n\nat a high level, you give forge something to run and it creates an isolated environment for it. something like:\n\n`forge run alpine:3.20 /bin/sh`\nand you get a shell running inside a container.\n\nbut forge isn't using docker underneath. i built the main pieces myself using linux primitives.\n\n**6**stages, each one a complete working system\n\n**12**internal packages, one per kernel concept\n\n**1**dependency outside the standard library (golang.org/x/sys)\n\n**0**shell-outs to docker, runc, ip or iptables\n\nthe goal was never to build something that replaces docker. i wanted to understand what a container actually is and what has to happen underneath that one simple command.\n\n## but what actually is a container?\n\nthis was probably the biggest thing i wanted to understand.\n\nwhen you hear \"container\", it can sound like you're starting some kind of tiny virtual computer. you're not.\n\nat the end of the day, a container is still a normal process running on your computer. the difference is that the process is given boundaries.\n\nit can be made to see its own processes, have its own hostname, have its own filesystem and network, and have limits on how many resources it is allowed to use.\n\nso the container feels like its own little machine, even though it's actually running on the same linux kernel as everything else.\n\nthe idea sounds pretty simple when you say it like that. actually making it work is a different story.\n\n## what forge does\n\ni built forge in six stages, with each stage adding another important part of what we normally think of as a container.\n\n### stage 1: process isolation\n\nthe first thing i needed was isolation.\n\nforge uses linux namespaces to give processes their own view of certain parts of the system. for example, a process inside the container can have its own pid namespace, which means it can see itself as pid 1 instead of seeing every process running on my laptop.\n\nit can also have its own hostname and mount namespace.\n\nthe part that actually asks the kernel for this is tiny. a `Config` says which namespaces you want, and that turns into flags for `clone(2)`:\n\n`internal/namespace/namespace.go`\n\n```\n// CloneFlags returns the clone(2) flags that create the requested namespaces.\n//\n// This is pure computation and is deliberately separated from Apply so the\n// mapping from Config to kernel flags is unit-testable without root.\nfunc (c Config) CloneFlags() uintptr {\n\tvar flags uintptr\n\tif c.PID {\n\t\tflags |= syscall.CLONE_NEWPID\n\t}\n\tif c.UTS {\n\t\tflags |= syscall.CLONE_NEWUTS\n\t}\n\tif c.Mount {\n\t\tflags |= syscall.CLONE_NEWNS\n\t}\n\tif c.Net {\n\t\tflags |= syscall.CLONE_NEWNET\n\t}\n\treturn flags\n}\n```\nthat was the first thing that surprised me. asking for isolation is one line per namespace. making the isolation actually hold is the rest of the project.\n\nthe very first example of that is the mount namespace. `CLONE_NEWNS` gives you a *copy* of the host's mount table, and the copy inherits each mount's propagation type. on any systemd host `/` is shared, so a mount made inside the container would still travel back out to the host. so the child has to detach the tree itself, from inside:\n\n```\nfunc makeMountTreePrivate() error {\n\tconst (\n\t\tsource = \"none\"\n\t\ttarget = \"/\"\n\t\tfstype = \"\"\n\t\tdata   = \"\"\n\t)\n\tif err := syscall.Mount(source, target, fstype, syscall.MS_REC|syscall.MS_PRIVATE, data); err != nil {\n\t\treturn fmt.Errorf(\"making mount tree private: %w\", translatePermission(err))\n\t}\n\treturn nil\n}\n```\na new mount namespace is not an empty one. it starts as a copy of its parent's, propagation types and all. without the recursive `MS_PRIVATE` above, every mount the container makes would still show up on the host. the namespace would exist, and the isolation still wouldn't.\n\nthis was the first point where forge started feeling like an actual container instead of just another go program starting a process.\n\n### stage 2: filesystem isolation\n\nnext was the filesystem.\n\nif i start a container, i don't want the process to simply have access to my entire computer.\n\nforge creates an isolated filesystem environment for the container and uses things like mounts and `pivot_root` to give the process its own view of the filesystem.\n\nso when you're inside the container, `/` is no longer simply the `/` from my laptop. it's the container's root filesystem.\n\n`internal/mount/apply.go`\n\n```\n// PivotRoot makes newRoot the calling process's root filesystem and detaches\n// the old one.\nfunc PivotRoot(newRoot string) error {\n\tnewRoot = filepath.Clean(newRoot)\n\t// The kernel refuses a new root that is not a mount point, with a bare\n\t// EINVAL that explains nothing. Say what it means instead.\n\tmounted, err := IsMountPoint(newRoot)\n\tif err != nil {\n\t\treturn err\n\t}\n\tif !mounted {\n\t\treturn fmt.Errorf(\"%w: %q is a plain directory; bind it onto itself first\", ErrRootNotMountPoint, newRoot)\n\t}\n\tputOld := filepath.Join(newRoot, oldRootDirName)\n\tif err := os.Mkdir(putOld, oldRootPerm); err != nil && !os.IsExist(err) {\n\t\treturn fmt.Errorf(\"creating %q for pivot_root: %w\", putOld, err)\n\t}\n\t// Entering the new root first means the process holds no reference to the\n\t// old one when it is detached below.\n\tif err := unix.Chdir(newRoot); err != nil {\n\t\treturn fmt.Errorf(\"entering the new root %q: %w\", newRoot, err)\n\t}\n\tif err := unix.PivotRoot(newRoot, putOld); err != nil {\n\t\treturn fmt.Errorf(\"pivot_root to %q: %w\", newRoot, translatePermission(err))\n\t}\n\t// \"/\" now means the new root, and the old one hangs off it.\n\tif err := unix.Chdir(\"/\"); err != nil {\n\t\treturn fmt.Errorf(\"entering the pivoted root: %w\", err)\n\t}\n\toldRoot := string(filepath.Separator) + oldRootDirName\n\tif err := unix.Unmount(oldRoot, unix.MNT_DETACH); err != nil {\n\t\treturn fmt.Errorf(\"detaching the old root at %q: %w\", oldRoot, translatePermission(err))\n\t}\n\t// ... the empty directory the old root hung from is removed here.\n\treturn nil\n}\n```\n`chroot` only changes the calling process's root *directory*. the old root stays mounted, stays listed in the process's own mount table, and is still reachable through the classic `mkdir tmp; chroot tmp; chdir(../../..)` walk. after `pivot_root` and the detach above, there is nothing left to walk back to.\n\nthis was also where i started seeing how many little details are involved in something that looks extremely simple from the outside.\n\none of those details is that a mount destination is a path *inside* a root filesystem i didn't create. resolving `/etc/hosts` the way the host would is exactly how a bind mount ends up writing to the host's `/etc`, so forge resolves every path component itself, rebasing absolute symlinks against the container's root instead of following them out.\n\n### stage 3: resource limits\n\nisolation isn't enough.\n\nwhat happens if a process inside the container decides to use as much memory as possible? or creates thousands of processes? or tries to use all of the cpu?\n\nthat's where cgroups come in.\n\nforge uses cgroups to put limits on resources such as memory, cpu and the number of processes.\n\nwith cgroups v2 a \"limit\" is genuinely just a string written into a file, so the whole thing splits cleanly into a pure function that decides what the kernel is told, and the write itself:\n\n`internal/cgroup/cgroup.go`\n\n```\nfunc (l Limits) Files() []File {\n\tvar files []File\n\tif l.MemoryMax != nil {\n\t\tfiles = append(files,\n\t\t\tFile{Name: \"memory.max\", Value: l.MemoryMax.String()},\n\t\t\t// Swap is limited alongside memory, never independently.\n\t\t\tFile{Name: \"memory.swap.max\", Value: swapFor(*l.MemoryMax), Optional: true},\n\t\t)\n\t}\n\tif l.CPU != nil {\n\t\tfiles = append(files, File{Name: \"cpu.max\", Value: l.CPU.String()})\n\t}\n\tif l.CPUWeight != nil {\n\t\tfiles = append(files, File{Name: \"cpu.weight\", Value: l.CPUWeight.String()})\n\t}\n\tif l.PIDsMax != nil {\n\t\tvalue := Unlimited\n\t\tif *l.PIDsMax >= 0 {\n\t\t\tvalue = strconv.FormatInt(*l.PIDsMax, 10)\n\t\t}\n\t\tfiles = append(files, File{Name: \"pids.max\", Value: value})\n\t}\n\treturn files\n}\n```\nand joining the cgroup is one write too:\n\n`internal/cgroup/apply.go`\n\n```\n// addProc makes the process a member of the cgroup at dir by writing its PID\n// to cgroup.procs.\n//\n// Writing a PID moves the whole thread group, and every process it forks from\n// then on is a member too. It does not move processes it has *already* forked,\n// which is why internal/runtime attaches the container's init before the init\n// is allowed to proceed past its handshake.\nfunc addProc(dir string, pid int) error {\n\treturn writeControlFile(dir, fileProcs, strconv.Itoa(pid))\n}\n```\nthat comment is the whole timing problem in stage 3. limits have to be in place *before* the container runs its first instruction, otherwise there is a window where an unlimited process already exists.\n\nso the container isn't just isolated. it also has limits on what it is allowed to consume.\n\n### stage 4: networking\n\nthen came networking.\n\na container isn't very useful if it can't communicate with anything.\n\nforge creates isolated network namespaces and sets up networking using things like veth pairs and a bridge. it also handles ip allocation and nat so that containers can communicate and reach the outside world.\n\nthis is the part i underestimated the most. everything the kernel is told here goes over a raw netlink socket, with no netlink library and no shelling out to `ip` or `iptables`. which means a netlink message is just a byte layout you have to get exactly right:\n\n`internal/network/network.go`\n\n```\n// nlAttr encodes one netlink attribute: a length, a type, the payload, and\n// enough padding to align the next attribute.\n//\n// The encoded length covers the header and the payload but *not* the padding,\n// which is why this cannot be a simple append of a header to a body.\nfunc nlAttr(typ uint16, payload []byte) []byte {\n\tlength := nlAttrHdrLen + len(payload)\n\tbuf := make([]byte, nlAlign(length))\n\torder.PutUint16(buf[0:2], uint16(length))\n\torder.PutUint16(buf[2:4], typ)\n\tcopy(buf[nlAttrHdrLen:], payload)\n\treturn buf\n}\n```\nand once you can build attributes, an operation that sounds enormous (\"move this interface into the container's network namespace\") turns out to be one message:\n\n`internal/network/namespace.go`\n\n```\nfunc moveLinkToNetns(c *nlConn, index int32, pid int) error {\n\tbody := concat(\n\t\tifInfoMsg(index, 0, 0),\n\t\tnlAttrU32(unix.IFLA_NET_NS_PID, uint32(pid)),\n\t)\n\tif err := c.execute(unix.RTM_NEWLINK, 0, body); err != nil {\n\t\treturn fmt.Errorf(\"moving interface %d into the network namespace of pid %d: %w\", index, pid, err)\n\t}\n\treturn nil\n}\n```\nforge never enters the container's network namespace to configure it. `setns(2)` changes the namespace of the calling *thread*, and in go that means locking an os thread and hoping nothing migrates. get it wrong and forge itself is left sitting inside a container's network namespace. so the parent does the one thing only it can do, which is move the interface into a namespace it can name by pid, and the container configures its own interface from a plain description that arrived over a pipe.\n\nthis was probably one of the parts where the command:\n\n`forge run ...`\nhides the most work.\n\nbehind that one command, forge has to create and connect a whole network environment for the container.\n\n### stage 5: images\n\nat this point i could create an isolated environment, but what exactly am i going to run inside it?\n\nthat's where container images come in.\n\nforge supports oci images. so when i run something like:\n\n`forge run alpine:3.20 /bin/sh`\nforge can pull the image, verify it, unpack its layers, and use that filesystem as the container's root filesystem.\n\nverification is the part i cared about most here. every document and every blob is named by the sha-256 of its own bytes, so it can be checked at every boundary the bytes cross, and forge checks at all of them: in flight as the registry streams them, again on the write to the cache, and again when a layer is decompressed for use.\n\n`internal/image/blob.go`\n\n```\n// FetchBlob streams one blob into w, verifying it as the bytes pass (FR-5.2).\n//\n// Both the digest and the length are checked. A hash mismatch alone would catch\n// a truncated response, but \"the registry sent 3 MB of the 5 MB it promised\" is\n// the sentence an operator can act on, and \"the hash was wrong\" is not.\nfunc (c *Client) FetchBlob(ctx context.Context, ref Reference, d Descriptor, w io.Writer) error {\n\thasher, err := newHasher(d.Digest)\n\tif err != nil {\n\t\treturn err\n\t}\n\tresp, err := c.get(ctx, ref, c.endpoint(ref, \"blobs\", d.Digest), nil)\n\tif err != nil {\n\t\treturn err\n\t}\n\tdefer drain(resp, c.logger)\n\twritten, err := io.Copy(io.MultiWriter(w, hasher), resp.Body)\n\tif err != nil {\n\t\tif ctxErr := ctx.Err(); ctxErr != nil {\n\t\t\treturn ctxErr\n\t\t}\n\t\treturn fmt.Errorf(\"%w: downloading %s after %d bytes: %w\",\n\t\t\tErrRegistryUnavailable, d.Digest, written, err)\n\t}\n\tif d.Size > 0 && written != d.Size {\n\t\treturn fmt.Errorf(\"%w: %s sent %d bytes for %s, but its descriptor says %d\",\n\t\t\tErrDigestMismatch, ref.Host(), written, d.Digest, d.Size)\n\t}\n\tif computed := formatDigest(hasher); computed != d.Digest {\n\t\treturn fmt.Errorf(\"%w: %d bytes from %s hash to %s, but were requested as %s\",\n\t\t\tErrDigestMismatch, written, ref.Host(), computed, d.Digest)\n\t}\n\treturn nil\n}\n```\nit also has image caching so that already downloaded content doesn't have to be downloaded again every time. layers are keyed by digest, so the second run of the same image downloads nothing.\n\nthis is integrity, not authenticity. it proves the bytes are the bytes that digest names. it does not prove who made them. signature and provenance verification is out of scope for forge.\n\nthis was the part that made the project feel much closer to an actual container runtime.\n\n### stage 6: the runtime\n\nthe final stage was taking all of these pieces and turning them into an actual runtime.\n\ninstead of having a collection of things that can create namespaces or configure networking, forge now has a container lifecycle.\n\nyou can:\n\n**run**create and start a container, attached\n\n**ps**what is running, from the on-disk state store\n\n**exec**a second process inside its namespaces\n\n**logs**captured stdout/stderr, attached or not\n\n**stop**sigterm, then sigkill after the timeout\n\n**rm**remove it and everything still held for it\n\nyou can start a container, see which containers are running, execute a command inside an existing container, read its logs, stop it and finally remove it.\n\n**sudo forge run --keep alpine:3.20 /bin/sh -c 'while :; do date; sleep 1; done'**\n*# in another terminal*\n**sudo forge ps**\nCONTAINER ID   IMAGE          COMMAND        STATUS    CREATED         PID\n7f3c9a1b2d04   alpine:3.20    /bin/sh -c …   running   12 seconds ago  48213\n**sudo forge exec 7f3c9a1b2d04 /bin/ps**\n*# the container's processes, not the host's*\nthat's when forge finally felt like a complete project to me.\n\n### the whole thing, in one picture\n\nby stage 6 there are six primitives that all have to happen in a particular order, split across two processes, because most of what makes a container a container can only be done by code already running *inside* the new namespaces. so forge doesn't start the container's binary directly. it starts *itself* again, and that second copy does the rest to itself.\n\n`1`resolve the image reference (pure, no i/o)\n\n`2`fetch the manifest, resolving the tag to an immutable digest\n\n`3`download only the layers the blob cache is missing\n\n`4`verify digests: on the wire, on the write, and at use\n\n`5`construct the rootfs: layers unpacked, base first\n\n`6`build the mount plan its init will apply\n\n`7`create the cgroup leaf and write the limits\n\n`8`bridge, nat and an address, claimed before there is a container to leak it\n\n`9`start the child, attach its pid to the cgroup, move the interface into its namespace, plug the host end into the bridge\n\n`10`send the payload down the pipe, releasing the child\n\n`11`wait: supervise until the container exits\n\n`12`unwind the cleanup stack, in reverse\n\n`·`blocked on the payload pipe, running nothing\n\n`a`\n\n`namespace.Apply`: mount tree made private first, or every mount below it propagates to the host\n`b`\n\n`network.Configure`: the pushed-in interface brought up and routed, while netlink is all that is needed\n`c`\n\n`mount.Apply`: made while the host filesystem is still reachable, because bind sources are host paths\n`d`\n\n`mount.PivotRoot`: \"/\" becomes the container's root, the old one is detached\n`e`\n\n`chdir`: the working directory, now a container path\n`f`\n\n`execve`: the process becomes the container\nthe child side of that is small enough to read in one go, and the order is load-bearing at every single line:\n\n`internal/runtime/init.go`\n\n```\n// Init is the container's entry point, executed by the re-exec'd forge binary\n// inside the new namespaces created by clone(2).\nfunc Init() error {\n\tpayload, err := readInitPayload()\n\tif err != nil {\n\t\treturn err\n\t}\n\tif err := namespace.Apply(payload.Namespace); err != nil {\n\t\treturn err\n\t}\n\tif err := configureNetwork(payload); err != nil {\n\t\treturn err\n\t}\n\tif payload.Mount != nil {\n\t\tif err := mount.Apply(*payload.Mount); err != nil {\n\t\t\treturn err\n\t\t}\n\t\tif err := mount.PivotRoot(payload.Mount.Root); err != nil {\n\t\t\treturn err\n\t\t}\n\t}\n\tif err := enterWorkingDir(payload.WorkingDir); err != nil {\n\t\treturn err\n\t}\n\tpath, err := resolveCommand(payload.Command[0], payload.Env)\n\tif err != nil {\n\t\treturn err\n\t}\n\tif err := syscall.Exec(path, payload.Command, payload.Env); err != nil {\n\t\treturn fmt.Errorf(\"executing %s: %w\", path, err)\n\t}\n\t// Unreachable: a successful execve never returns.\n\treturn errors.New(\"execve returned without an error\")\n}\n```\non success this function never returns. `execve` replaces the process with the container's binary, which inherits its pid. and inside a pid namespace, that pid is 1.\n\n## why didn't i just use docker?\n\nbecause that wasn't the point.\n\nif i just wanted to run containers, i would use docker. the whole reason i made forge was to understand what docker and other container runtimes are actually doing underneath.\n\nusing something is one thing. building a smaller version yourself is a completely different experience.\n\nwhen you're forced to implement the pieces yourself, you can't just say \"docker handles that.\"\n\nyou have to figure out what \"that\" actually means.\n\nhow does the process get isolated? how does the filesystem change? how does networking get connected? how are resources limited? how does the image become a filesystem? how does everything get cleaned up when the container stops?\n\nthose were the things i wanted answers to.\n\n## what i learned from building it\n\nthe biggest thing i learned is that a container isn't really one big complicated thing. it's a bunch of smaller linux features working together.\n\nnamespaces handle isolation. cgroups handle resource limits. the filesystem gives the process its own environment. networking connects it to the outside world. oci images provide the filesystem and configuration. and the runtime puts all of those pieces together.\n\nonce you break it down like that, the word \"container\" becomes a lot less mysterious.\n\n## and obviously, things went wrong\n\nthis wasn't just me writing some code and everything magically working.\n\nthere were plenty of moments where something looked correct but didn't actually work the way i expected, especially with things like networking, cleanup, process handling and filesystem setup.\n\nand that's actually one of the reasons i wanted to build this instead of just reading about containers.\n\nwhen something breaks, you can't skip over the details. you have to figure out exactly which part of the system is responsible.\n\nthat ended up teaching me a lot more than just reading a diagram of how containers work.\n\n## how i tested it\n\ni also didn't want forge to be one of those projects where the readme says it works because one command worked once.\n\nthe project has unit tests and integration tests covering the different parts of the runtime. i also use race detection and linting as part of the validation.\n\n```\nmake test               # unit tests, with -race, no root required\nmake test-integration   # privileged integration tests (root, linux)\nmake lint               # golangci-lint\n```\nthe split matters more than it looks. everything that is pure (clone flags, cgroup limit files, netlink byte layouts, subnet arithmetic, mount-path resolution) is a function from values to bytes, so it can be asserted without a kernel. everything that actually touches the kernel lives behind a build tag and runs as root.\n\nthe final testing isn't just about checking whether a container starts. i need to know that processes are actually isolated, the filesystem is actually isolated, resource limits actually work, networking actually works, images are handled correctly, commands like `exec`, `logs`, `stop` and `rm` work, multiple containers can coexist, and cleanup actually happens.\n\nbecause a container that starts but leaves half its resources behind isn't really a finished runtime.\n\n## forge isn't docker\n\njust to be clear, forge isn't trying to be a replacement for docker or any other production container runtime.\n\nit's a learning project.\n\ni wanted something small enough that i could actually understand the whole thing.\n\nsomething where i could go from:\n\n**process**still just a normal process on the host\n\n**namespace**its own pids, hostname, mounts, network\n\n**filesystem**pivot_root, bind mounts, its own \"/\"\n\n**cgroup**memory, cpu and pid ceilings\n\n**network**veth, bridge, an address, nat\n\n**image**oci layers, verified and unpacked\n\n**runtime**a lifecycle that ties all six together\n\nand understand what each part is doing.\n\nthat's what forge is.\n\nforge is an educational systems project. it is not production-hardened and should not be used to run untrusted workloads. there is no seccomp, no apparmor or selinux integration, no rootless mode, no image building and no orchestration. those are deliberate non-goals, not a todo list.\n\n## if you're not a tech person\n\nif everything above sounded confusing, here's the easiest way i can explain it.\n\nso basically imagine you have a game on your laptop. it works perfectly. you send that same game to your friend, and they try to run it but it doesn't work.\n\nmaybe their laptop has different software installed. maybe they're missing something your laptop has. maybe their settings are different.\n\nthat's the classic:\n\n**\"it works on my laptop.\"**\n\na container is kind of like giving that game its own little room with the things it needs inside.\n\ninstead of depending completely on what the rest of the laptop looks like, it gets its own controlled environment.\n\n**forge is something i built to create those little rooms.**\n\nthat's probably the simplest way i can explain the entire project.\n\nthanks for reading.\n\n### Related Project\n\n#### Forge\n\nA container runtime built from scratch in Go using Linux primitives and raw kernel interfaces.","body_html":"<h1 id=\"inside-forge-writing-a-container-runtime-from-scratch\">Inside Forge | Writing a Container Runtime From Scratch</h1>\n<h1 id=\"i-built-my-own-container-runtime-in-go\">i built my own container runtime in go</h1>\n<p>i&#39;ve used containers before. you run something, it works, and you don&#39;t really think about what is happening underneath. you just know that somehow this thing called a &quot;container&quot; makes your application run in its own environment.</p>\n<p>and honestly, i didn&#39;t really like that. i wanted to know what is actually happening when you run a container. what makes it isolated? how does it get its own filesystem? how does it get its own network? how does it stop a process from using all the memory on the machine?</p>\n<p>so instead of just reading about it, i decided to build one myself.</p>\n<p>that&#39;s how <strong>forge</strong> started.</p>\n<h2 id=\"so-what-is-forge\">so, what is forge?</h2>\n<p>forge is a small container runtime that i built from scratch in go. the source is on github ↗</p>\n<p>at a high level, you give forge something to run and it creates an isolated environment for it. something like:</p>\n<p><code>forge run alpine:3.20 /bin/sh</code>\nand you get a shell running inside a container.</p>\n<p>but forge isn&#39;t using docker underneath. i built the main pieces myself using linux primitives.</p>\n<p>*<em>6*</em>stages, each one a complete working system</p>\n<p><strong>12</strong>internal packages, one per kernel concept</p>\n<p>*<em>1*</em>dependency outside the standard library (golang.org/x/sys)</p>\n<p>*<em>0*</em>shell-outs to docker, runc, ip or iptables</p>\n<p>the goal was never to build something that replaces docker. i wanted to understand what a container actually is and what has to happen underneath that one simple command.</p>\n<h2 id=\"but-what-actually-is-a-container\">but what actually is a container?</h2>\n<p>this was probably the biggest thing i wanted to understand.</p>\n<p>when you hear &quot;container&quot;, it can sound like you&#39;re starting some kind of tiny virtual computer. you&#39;re not.</p>\n<p>at the end of the day, a container is still a normal process running on your computer. the difference is that the process is given boundaries.</p>\n<p>it can be made to see its own processes, have its own hostname, have its own filesystem and network, and have limits on how many resources it is allowed to use.</p>\n<p>so the container feels like its own little machine, even though it&#39;s actually running on the same linux kernel as everything else.</p>\n<p>the idea sounds pretty simple when you say it like that. actually making it work is a different story.</p>\n<h2 id=\"what-forge-does\">what forge does</h2>\n<p>i built forge in six stages, with each stage adding another important part of what we normally think of as a container.</p>\n<h3 id=\"stage-1-process-isolation\">stage 1: process isolation</h3>\n<p>the first thing i needed was isolation.</p>\n<p>forge uses linux namespaces to give processes their own view of certain parts of the system. for example, a process inside the container can have its own pid namespace, which means it can see itself as pid 1 instead of seeing every process running on my laptop.</p>\n<p>it can also have its own hostname and mount namespace.</p>\n<p>the part that actually asks the kernel for this is tiny. a <code>Config</code> says which namespaces you want, and that turns into flags for <code>clone(2)</code>:</p>\n<p><code>internal/namespace/namespace.go</code></p>\n<pre><code>// CloneFlags returns the clone(2) flags that create the requested namespaces.\n//\n// This is pure computation and is deliberately separated from Apply so the\n// mapping from Config to kernel flags is unit-testable without root.\nfunc (c Config) CloneFlags() uintptr {\n    var flags uintptr\n    if c.PID {\n        flags |= syscall.CLONE_NEWPID\n    }\n    if c.UTS {\n        flags |= syscall.CLONE_NEWUTS\n    }\n    if c.Mount {\n        flags |= syscall.CLONE_NEWNS\n    }\n    if c.Net {\n        flags |= syscall.CLONE_NEWNET\n    }\n    return flags\n}</code></pre>\n<p>that was the first thing that surprised me. asking for isolation is one line per namespace. making the isolation actually hold is the rest of the project.</p>\n<p>the very first example of that is the mount namespace. <code>CLONE_NEWNS</code> gives you a <em>copy</em> of the host&#39;s mount table, and the copy inherits each mount&#39;s propagation type. on any systemd host <code>/</code> is shared, so a mount made inside the container would still travel back out to the host. so the child has to detach the tree itself, from inside:</p>\n<pre><code>func makeMountTreePrivate() error {\n    const (\n        source = &quot;none&quot;\n        target = &quot;/&quot;\n        fstype = &quot;&quot;\n        data   = &quot;&quot;\n    )\n    if err := syscall.Mount(source, target, fstype, syscall.MS_REC|syscall.MS_PRIVATE, data); err != nil {\n        return fmt.Errorf(&quot;making mount tree private: %w&quot;, translatePermission(err))\n    }\n    return nil\n}</code></pre>\n<p>a new mount namespace is not an empty one. it starts as a copy of its parent&#39;s, propagation types and all. without the recursive <code>MS_PRIVATE</code> above, every mount the container makes would still show up on the host. the namespace would exist, and the isolation still wouldn&#39;t.</p>\n<p>this was the first point where forge started feeling like an actual container instead of just another go program starting a process.</p>\n<h3 id=\"stage-2-filesystem-isolation\">stage 2: filesystem isolation</h3>\n<p>next was the filesystem.</p>\n<p>if i start a container, i don&#39;t want the process to simply have access to my entire computer.</p>\n<p>forge creates an isolated filesystem environment for the container and uses things like mounts and <code>pivot_root</code> to give the process its own view of the filesystem.</p>\n<p>so when you&#39;re inside the container, <code>/</code> is no longer simply the <code>/</code> from my laptop. it&#39;s the container&#39;s root filesystem.</p>\n<p><code>internal/mount/apply.go</code></p>\n<pre><code>// PivotRoot makes newRoot the calling process&#39;s root filesystem and detaches\n// the old one.\nfunc PivotRoot(newRoot string) error {\n    newRoot = filepath.Clean(newRoot)\n    // The kernel refuses a new root that is not a mount point, with a bare\n    // EINVAL that explains nothing. Say what it means instead.\n    mounted, err := IsMountPoint(newRoot)\n    if err != nil {\n        return err\n    }\n    if !mounted {\n        return fmt.Errorf(&quot;%w: %q is a plain directory; bind it onto itself first&quot;, ErrRootNotMountPoint, newRoot)\n    }\n    putOld := filepath.Join(newRoot, oldRootDirName)\n    if err := os.Mkdir(putOld, oldRootPerm); err != nil &amp;&amp; !os.IsExist(err) {\n        return fmt.Errorf(&quot;creating %q for pivot_root: %w&quot;, putOld, err)\n    }\n    // Entering the new root first means the process holds no reference to the\n    // old one when it is detached below.\n    if err := unix.Chdir(newRoot); err != nil {\n        return fmt.Errorf(&quot;entering the new root %q: %w&quot;, newRoot, err)\n    }\n    if err := unix.PivotRoot(newRoot, putOld); err != nil {\n        return fmt.Errorf(&quot;pivot_root to %q: %w&quot;, newRoot, translatePermission(err))\n    }\n    // &quot;/&quot; now means the new root, and the old one hangs off it.\n    if err := unix.Chdir(&quot;/&quot;); err != nil {\n        return fmt.Errorf(&quot;entering the pivoted root: %w&quot;, err)\n    }\n    oldRoot := string(filepath.Separator) + oldRootDirName\n    if err := unix.Unmount(oldRoot, unix.MNT_DETACH); err != nil {\n        return fmt.Errorf(&quot;detaching the old root at %q: %w&quot;, oldRoot, translatePermission(err))\n    }\n    // ... the empty directory the old root hung from is removed here.\n    return nil\n}</code></pre>\n<p><code>chroot</code> only changes the calling process&#39;s root <em>directory</em>. the old root stays mounted, stays listed in the process&#39;s own mount table, and is still reachable through the classic <code>mkdir tmp; chroot tmp; chdir(../../..)</code> walk. after <code>pivot_root</code> and the detach above, there is nothing left to walk back to.</p>\n<p>this was also where i started seeing how many little details are involved in something that looks extremely simple from the outside.</p>\n<p>one of those details is that a mount destination is a path <em>inside</em> a root filesystem i didn&#39;t create. resolving <code>/etc/hosts</code> the way the host would is exactly how a bind mount ends up writing to the host&#39;s <code>/etc</code>, so forge resolves every path component itself, rebasing absolute symlinks against the container&#39;s root instead of following them out.</p>\n<h3 id=\"stage-3-resource-limits\">stage 3: resource limits</h3>\n<p>isolation isn&#39;t enough.</p>\n<p>what happens if a process inside the container decides to use as much memory as possible? or creates thousands of processes? or tries to use all of the cpu?</p>\n<p>that&#39;s where cgroups come in.</p>\n<p>forge uses cgroups to put limits on resources such as memory, cpu and the number of processes.</p>\n<p>with cgroups v2 a &quot;limit&quot; is genuinely just a string written into a file, so the whole thing splits cleanly into a pure function that decides what the kernel is told, and the write itself:</p>\n<p><code>internal/cgroup/cgroup.go</code></p>\n<pre><code>func (l Limits) Files() []File {\n    var files []File\n    if l.MemoryMax != nil {\n        files = append(files,\n            File{Name: &quot;memory.max&quot;, Value: l.MemoryMax.String()},\n            // Swap is limited alongside memory, never independently.\n            File{Name: &quot;memory.swap.max&quot;, Value: swapFor(*l.MemoryMax), Optional: true},\n        )\n    }\n    if l.CPU != nil {\n        files = append(files, File{Name: &quot;cpu.max&quot;, Value: l.CPU.String()})\n    }\n    if l.CPUWeight != nil {\n        files = append(files, File{Name: &quot;cpu.weight&quot;, Value: l.CPUWeight.String()})\n    }\n    if l.PIDsMax != nil {\n        value := Unlimited\n        if *l.PIDsMax &gt;= 0 {\n            value = strconv.FormatInt(*l.PIDsMax, 10)\n        }\n        files = append(files, File{Name: &quot;pids.max&quot;, Value: value})\n    }\n    return files\n}</code></pre>\n<p>and joining the cgroup is one write too:</p>\n<p><code>internal/cgroup/apply.go</code></p>\n<pre><code>// addProc makes the process a member of the cgroup at dir by writing its PID\n// to cgroup.procs.\n//\n// Writing a PID moves the whole thread group, and every process it forks from\n// then on is a member too. It does not move processes it has *already* forked,\n// which is why internal/runtime attaches the container&#39;s init before the init\n// is allowed to proceed past its handshake.\nfunc addProc(dir string, pid int) error {\n    return writeControlFile(dir, fileProcs, strconv.Itoa(pid))\n}</code></pre>\n<p>that comment is the whole timing problem in stage 3. limits have to be in place <em>before</em> the container runs its first instruction, otherwise there is a window where an unlimited process already exists.</p>\n<p>so the container isn&#39;t just isolated. it also has limits on what it is allowed to consume.</p>\n<h3 id=\"stage-4-networking\">stage 4: networking</h3>\n<p>then came networking.</p>\n<p>a container isn&#39;t very useful if it can&#39;t communicate with anything.</p>\n<p>forge creates isolated network namespaces and sets up networking using things like veth pairs and a bridge. it also handles ip allocation and nat so that containers can communicate and reach the outside world.</p>\n<p>this is the part i underestimated the most. everything the kernel is told here goes over a raw netlink socket, with no netlink library and no shelling out to <code>ip</code> or <code>iptables</code>. which means a netlink message is just a byte layout you have to get exactly right:</p>\n<p><code>internal/network/network.go</code></p>\n<pre><code>// nlAttr encodes one netlink attribute: a length, a type, the payload, and\n// enough padding to align the next attribute.\n//\n// The encoded length covers the header and the payload but *not* the padding,\n// which is why this cannot be a simple append of a header to a body.\nfunc nlAttr(typ uint16, payload []byte) []byte {\n    length := nlAttrHdrLen + len(payload)\n    buf := make([]byte, nlAlign(length))\n    order.PutUint16(buf[0:2], uint16(length))\n    order.PutUint16(buf[2:4], typ)\n    copy(buf[nlAttrHdrLen:], payload)\n    return buf\n}</code></pre>\n<p>and once you can build attributes, an operation that sounds enormous (&quot;move this interface into the container&#39;s network namespace&quot;) turns out to be one message:</p>\n<p><code>internal/network/namespace.go</code></p>\n<pre><code>func moveLinkToNetns(c *nlConn, index int32, pid int) error {\n    body := concat(\n        ifInfoMsg(index, 0, 0),\n        nlAttrU32(unix.IFLA_NET_NS_PID, uint32(pid)),\n    )\n    if err := c.execute(unix.RTM_NEWLINK, 0, body); err != nil {\n        return fmt.Errorf(&quot;moving interface %d into the network namespace of pid %d: %w&quot;, index, pid, err)\n    }\n    return nil\n}</code></pre>\n<p>forge never enters the container&#39;s network namespace to configure it. <code>setns(2)</code> changes the namespace of the calling <em>thread</em>, and in go that means locking an os thread and hoping nothing migrates. get it wrong and forge itself is left sitting inside a container&#39;s network namespace. so the parent does the one thing only it can do, which is move the interface into a namespace it can name by pid, and the container configures its own interface from a plain description that arrived over a pipe.</p>\n<p>this was probably one of the parts where the command:</p>\n<p><code>forge run ...</code>\nhides the most work.</p>\n<p>behind that one command, forge has to create and connect a whole network environment for the container.</p>\n<h3 id=\"stage-5-images\">stage 5: images</h3>\n<p>at this point i could create an isolated environment, but what exactly am i going to run inside it?</p>\n<p>that&#39;s where container images come in.</p>\n<p>forge supports oci images. so when i run something like:</p>\n<p><code>forge run alpine:3.20 /bin/sh</code>\nforge can pull the image, verify it, unpack its layers, and use that filesystem as the container&#39;s root filesystem.</p>\n<p>verification is the part i cared about most here. every document and every blob is named by the sha-256 of its own bytes, so it can be checked at every boundary the bytes cross, and forge checks at all of them: in flight as the registry streams them, again on the write to the cache, and again when a layer is decompressed for use.</p>\n<p><code>internal/image/blob.go</code></p>\n<pre><code>// FetchBlob streams one blob into w, verifying it as the bytes pass (FR-5.2).\n//\n// Both the digest and the length are checked. A hash mismatch alone would catch\n// a truncated response, but &quot;the registry sent 3 MB of the 5 MB it promised&quot; is\n// the sentence an operator can act on, and &quot;the hash was wrong&quot; is not.\nfunc (c *Client) FetchBlob(ctx context.Context, ref Reference, d Descriptor, w io.Writer) error {\n    hasher, err := newHasher(d.Digest)\n    if err != nil {\n        return err\n    }\n    resp, err := c.get(ctx, ref, c.endpoint(ref, &quot;blobs&quot;, d.Digest), nil)\n    if err != nil {\n        return err\n    }\n    defer drain(resp, c.logger)\n    written, err := io.Copy(io.MultiWriter(w, hasher), resp.Body)\n    if err != nil {\n        if ctxErr := ctx.Err(); ctxErr != nil {\n            return ctxErr\n        }\n        return fmt.Errorf(&quot;%w: downloading %s after %d bytes: %w&quot;,\n            ErrRegistryUnavailable, d.Digest, written, err)\n    }\n    if d.Size &gt; 0 &amp;&amp; written != d.Size {\n        return fmt.Errorf(&quot;%w: %s sent %d bytes for %s, but its descriptor says %d&quot;,\n            ErrDigestMismatch, ref.Host(), written, d.Digest, d.Size)\n    }\n    if computed := formatDigest(hasher); computed != d.Digest {\n        return fmt.Errorf(&quot;%w: %d bytes from %s hash to %s, but were requested as %s&quot;,\n            ErrDigestMismatch, written, ref.Host(), computed, d.Digest)\n    }\n    return nil\n}</code></pre>\n<p>it also has image caching so that already downloaded content doesn&#39;t have to be downloaded again every time. layers are keyed by digest, so the second run of the same image downloads nothing.</p>\n<p>this is integrity, not authenticity. it proves the bytes are the bytes that digest names. it does not prove who made them. signature and provenance verification is out of scope for forge.</p>\n<p>this was the part that made the project feel much closer to an actual container runtime.</p>\n<h3 id=\"stage-6-the-runtime\">stage 6: the runtime</h3>\n<p>the final stage was taking all of these pieces and turning them into an actual runtime.</p>\n<p>instead of having a collection of things that can create namespaces or configure networking, forge now has a container lifecycle.</p>\n<p>you can:</p>\n<p><strong>run</strong>create and start a container, attached</p>\n<p><strong>ps</strong>what is running, from the on-disk state store</p>\n<p><strong>exec</strong>a second process inside its namespaces</p>\n<p><strong>logs</strong>captured stdout/stderr, attached or not</p>\n<p><strong>stop</strong>sigterm, then sigkill after the timeout</p>\n<p><strong>rm</strong>remove it and everything still held for it</p>\n<p>you can start a container, see which containers are running, execute a command inside an existing container, read its logs, stop it and finally remove it.</p>\n<p><strong>sudo forge run --keep alpine:3.20 /bin/sh -c &#39;while :; do date; sleep 1; done&#39;</strong>\n<em># in another terminal</em>\n<strong>sudo forge ps</strong>\nCONTAINER ID   IMAGE          COMMAND        STATUS    CREATED         PID\n7f3c9a1b2d04   alpine:3.20    /bin/sh -c …   running   12 seconds ago  48213\n<strong>sudo forge exec 7f3c9a1b2d04 /bin/ps</strong>\n<em># the container&#39;s processes, not the host&#39;s</em>\nthat&#39;s when forge finally felt like a complete project to me.</p>\n<h3 id=\"the-whole-thing-in-one-picture\">the whole thing, in one picture</h3>\n<p>by stage 6 there are six primitives that all have to happen in a particular order, split across two processes, because most of what makes a container a container can only be done by code already running <em>inside</em> the new namespaces. so forge doesn&#39;t start the container&#39;s binary directly. it starts <em>itself</em> again, and that second copy does the rest to itself.</p>\n<p><code>1</code>resolve the image reference (pure, no i/o)</p>\n<p><code>2</code>fetch the manifest, resolving the tag to an immutable digest</p>\n<p><code>3</code>download only the layers the blob cache is missing</p>\n<p><code>4</code>verify digests: on the wire, on the write, and at use</p>\n<p><code>5</code>construct the rootfs: layers unpacked, base first</p>\n<p><code>6</code>build the mount plan its init will apply</p>\n<p><code>7</code>create the cgroup leaf and write the limits</p>\n<p><code>8</code>bridge, nat and an address, claimed before there is a container to leak it</p>\n<p><code>9</code>start the child, attach its pid to the cgroup, move the interface into its namespace, plug the host end into the bridge</p>\n<p><code>10</code>send the payload down the pipe, releasing the child</p>\n<p><code>11</code>wait: supervise until the container exits</p>\n<p><code>12</code>unwind the cleanup stack, in reverse</p>\n<p><code>·</code>blocked on the payload pipe, running nothing</p>\n<p><code>a</code></p>\n<p><code>namespace.Apply</code>: mount tree made private first, or every mount below it propagates to the host\n<code>b</code></p>\n<p><code>network.Configure</code>: the pushed-in interface brought up and routed, while netlink is all that is needed\n<code>c</code></p>\n<p><code>mount.Apply</code>: made while the host filesystem is still reachable, because bind sources are host paths\n<code>d</code></p>\n<p><code>mount.PivotRoot</code>: &quot;/&quot; becomes the container&#39;s root, the old one is detached\n<code>e</code></p>\n<p><code>chdir</code>: the working directory, now a container path\n<code>f</code></p>\n<p><code>execve</code>: the process becomes the container\nthe child side of that is small enough to read in one go, and the order is load-bearing at every single line:</p>\n<p><code>internal/runtime/init.go</code></p>\n<pre><code>// Init is the container&#39;s entry point, executed by the re-exec&#39;d forge binary\n// inside the new namespaces created by clone(2).\nfunc Init() error {\n    payload, err := readInitPayload()\n    if err != nil {\n        return err\n    }\n    if err := namespace.Apply(payload.Namespace); err != nil {\n        return err\n    }\n    if err := configureNetwork(payload); err != nil {\n        return err\n    }\n    if payload.Mount != nil {\n        if err := mount.Apply(*payload.Mount); err != nil {\n            return err\n        }\n        if err := mount.PivotRoot(payload.Mount.Root); err != nil {\n            return err\n        }\n    }\n    if err := enterWorkingDir(payload.WorkingDir); err != nil {\n        return err\n    }\n    path, err := resolveCommand(payload.Command[0], payload.Env)\n    if err != nil {\n        return err\n    }\n    if err := syscall.Exec(path, payload.Command, payload.Env); err != nil {\n        return fmt.Errorf(&quot;executing %s: %w&quot;, path, err)\n    }\n    // Unreachable: a successful execve never returns.\n    return errors.New(&quot;execve returned without an error&quot;)\n}</code></pre>\n<p>on success this function never returns. <code>execve</code> replaces the process with the container&#39;s binary, which inherits its pid. and inside a pid namespace, that pid is 1.</p>\n<h2 id=\"why-didn-t-i-just-use-docker\">why didn&#39;t i just use docker?</h2>\n<p>because that wasn&#39;t the point.</p>\n<p>if i just wanted to run containers, i would use docker. the whole reason i made forge was to understand what docker and other container runtimes are actually doing underneath.</p>\n<p>using something is one thing. building a smaller version yourself is a completely different experience.</p>\n<p>when you&#39;re forced to implement the pieces yourself, you can&#39;t just say &quot;docker handles that.&quot;</p>\n<p>you have to figure out what &quot;that&quot; actually means.</p>\n<p>how does the process get isolated? how does the filesystem change? how does networking get connected? how are resources limited? how does the image become a filesystem? how does everything get cleaned up when the container stops?</p>\n<p>those were the things i wanted answers to.</p>\n<h2 id=\"what-i-learned-from-building-it\">what i learned from building it</h2>\n<p>the biggest thing i learned is that a container isn&#39;t really one big complicated thing. it&#39;s a bunch of smaller linux features working together.</p>\n<p>namespaces handle isolation. cgroups handle resource limits. the filesystem gives the process its own environment. networking connects it to the outside world. oci images provide the filesystem and configuration. and the runtime puts all of those pieces together.</p>\n<p>once you break it down like that, the word &quot;container&quot; becomes a lot less mysterious.</p>\n<h2 id=\"and-obviously-things-went-wrong\">and obviously, things went wrong</h2>\n<p>this wasn&#39;t just me writing some code and everything magically working.</p>\n<p>there were plenty of moments where something looked correct but didn&#39;t actually work the way i expected, especially with things like networking, cleanup, process handling and filesystem setup.</p>\n<p>and that&#39;s actually one of the reasons i wanted to build this instead of just reading about containers.</p>\n<p>when something breaks, you can&#39;t skip over the details. you have to figure out exactly which part of the system is responsible.</p>\n<p>that ended up teaching me a lot more than just reading a diagram of how containers work.</p>\n<h2 id=\"how-i-tested-it\">how i tested it</h2>\n<p>i also didn&#39;t want forge to be one of those projects where the readme says it works because one command worked once.</p>\n<p>the project has unit tests and integration tests covering the different parts of the runtime. i also use race detection and linting as part of the validation.</p>\n<pre><code>make test               # unit tests, with -race, no root required\nmake test-integration   # privileged integration tests (root, linux)\nmake lint               # golangci-lint</code></pre>\n<p>the split matters more than it looks. everything that is pure (clone flags, cgroup limit files, netlink byte layouts, subnet arithmetic, mount-path resolution) is a function from values to bytes, so it can be asserted without a kernel. everything that actually touches the kernel lives behind a build tag and runs as root.</p>\n<p>the final testing isn&#39;t just about checking whether a container starts. i need to know that processes are actually isolated, the filesystem is actually isolated, resource limits actually work, networking actually works, images are handled correctly, commands like <code>exec</code>, <code>logs</code>, <code>stop</code> and <code>rm</code> work, multiple containers can coexist, and cleanup actually happens.</p>\n<p>because a container that starts but leaves half its resources behind isn&#39;t really a finished runtime.</p>\n<h2 id=\"forge-isn-t-docker\">forge isn&#39;t docker</h2>\n<p>just to be clear, forge isn&#39;t trying to be a replacement for docker or any other production container runtime.</p>\n<p>it&#39;s a learning project.</p>\n<p>i wanted something small enough that i could actually understand the whole thing.</p>\n<p>something where i could go from:</p>\n<p><strong>process</strong>still just a normal process on the host</p>\n<p><strong>namespace</strong>its own pids, hostname, mounts, network</p>\n<p><strong>filesystem</strong>pivot_root, bind mounts, its own &quot;/&quot;</p>\n<p><strong>cgroup</strong>memory, cpu and pid ceilings</p>\n<p><strong>network</strong>veth, bridge, an address, nat</p>\n<p><strong>image</strong>oci layers, verified and unpacked</p>\n<p><strong>runtime</strong>a lifecycle that ties all six together</p>\n<p>and understand what each part is doing.</p>\n<p>that&#39;s what forge is.</p>\n<p>forge is an educational systems project. it is not production-hardened and should not be used to run untrusted workloads. there is no seccomp, no apparmor or selinux integration, no rootless mode, no image building and no orchestration. those are deliberate non-goals, not a todo list.</p>\n<h2 id=\"if-you-re-not-a-tech-person\">if you&#39;re not a tech person</h2>\n<p>if everything above sounded confusing, here&#39;s the easiest way i can explain it.</p>\n<p>so basically imagine you have a game on your laptop. it works perfectly. you send that same game to your friend, and they try to run it but it doesn&#39;t work.</p>\n<p>maybe their laptop has different software installed. maybe they&#39;re missing something your laptop has. maybe their settings are different.</p>\n<p>that&#39;s the classic:</p>\n<p><strong>&quot;it works on my laptop.&quot;</strong></p>\n<p>a container is kind of like giving that game its own little room with the things it needs inside.</p>\n<p>instead of depending completely on what the rest of the laptop looks like, it gets its own controlled environment.</p>\n<p><strong>forge is something i built to create those little rooms.</strong></p>\n<p>that&#39;s probably the simplest way i can explain the entire project.</p>\n<p>thanks for reading.</p>\n<h3 id=\"related-project\">Related Project</h3>\n<h4 id=\"forge\">Forge</h4>\n<p>A container runtime built from scratch in Go using Linux primitives and raw kernel interfaces.</p>","headings":[{"level":1,"text":"Inside Forge | Writing a Container Runtime From Scratch","id":"inside-forge-writing-a-container-runtime-from-scratch"},{"level":1,"text":"i built my own container runtime in go","id":"i-built-my-own-container-runtime-in-go"},{"level":2,"text":"so, what is forge?","id":"so-what-is-forge"},{"level":2,"text":"but what actually is a container?","id":"but-what-actually-is-a-container"},{"level":2,"text":"what forge does","id":"what-forge-does"},{"level":3,"text":"stage 1: process isolation","id":"stage-1-process-isolation"},{"level":3,"text":"stage 2: filesystem isolation","id":"stage-2-filesystem-isolation"},{"level":3,"text":"stage 3: resource limits","id":"stage-3-resource-limits"},{"level":3,"text":"stage 4: networking","id":"stage-4-networking"},{"level":3,"text":"stage 5: images","id":"stage-5-images"},{"level":3,"text":"stage 6: the runtime","id":"stage-6-the-runtime"},{"level":3,"text":"the whole thing, in one picture","id":"the-whole-thing-in-one-picture"},{"level":2,"text":"why didn't i just use docker?","id":"why-didn-t-i-just-use-docker"},{"level":2,"text":"what i learned from building it","id":"what-i-learned-from-building-it"},{"level":2,"text":"and obviously, things went wrong","id":"and-obviously-things-went-wrong"},{"level":2,"text":"how i tested it","id":"how-i-tested-it"},{"level":2,"text":"forge isn't docker","id":"forge-isn-t-docker"},{"level":2,"text":"if you're not a tech person","id":"if-you-re-not-a-tech-person"},{"level":3,"text":"Related Project","id":"related-project"}]}}