{"article":{"slug":"you-can-run-git-on-object-storage-if-you-re-make-packfiles","title":"You can run git on object storage if you re-make packfiles","subtitle":null,"summary":"Building ObjGit, Tigris explains why Git packfiles fight object storage, how remaking packfiles unlocks workable remote repositories, and what that means for Git servers backed by S3-style buckets.","content_type":"blog_post","language":"en","canonical_url":"https://www.tigrisdata.com/blog/objgit-packfiles/","author":{"name":"Xe Iaso","url":"https://xeiaso.net","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Tigris Data","url":"https://www.tigrisdata.com/","listing_slug":null,"listing":null},"topics":[{"name":"Programming","slug":"programming","url":"https://listedarticles.com/topics/programming"},{"name":"Open Source","slug":"open-source","url":"https://listedarticles.com/topics/open-source"},{"name":"Engineering","slug":"engineering","url":"https://listedarticles.com/topics/engineering"},{"name":"Performance","slug":"performance","url":"https://listedarticles.com/topics/performance"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":3408,"reading_minutes":15,"published_at":"2026-09-15T00:00:00.000Z","added_at":"2026-09-19T09:13:36.595Z","updated_at":"2026-09-19T09:13:36.595Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":false},"profile_url":"https://listedarticles.com/articles/you-can-run-git-on-object-storage-if-you-re-make-packfiles","markdown_url":"https://listedarticles.com/articles/you-can-run-git-on-object-storage-if-you-re-make-packfiles.md","example":false,"citation":"Xe Iaso, Tigris Data. \"You can run git on object storage if you re-make packfiles.\" 15 Sept 2026. https://www.tigrisdata.com/blog/objgit-packfiles/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://www.tigrisdata.com/blog/objgit-packfiles/"},"body_markdown":"# You can run git on object storage if you re-make packfiles\n\n[Blog](/blog/)\n\n# You can run git on object storage if you re-make packfiles\n\nGit packfiles were designed for mmap and local disk, so pulling one object out of a bucket means guessing at a byte range. I wrote a new format instead.\n\n## Contents\n\nIt sure seems that a bunch of companies are trying to ship a git product of some kind as of late. Wonder why that is.\n\nEither way, I’m building a Git server backed by object storage as an\n[open-source project](https://github.com/tigrisdata/objgit). It sounded simple\nenough to start: Git looks like a filesystem, so let’s use a filesystem as a\ntranslation layer on top of object storage to make Git speak object storage.\nThis model worked… [ok, I guess?](https://www.tigrisdata.com/blog/objgit/) But\nit didn’t work for real-world size repositories, so I needed a different\napproach. Git stores everything in\n[Objects](https://git-scm.com/book/en/v2/Git-Internals-Git-Objects), so why not\nstore those as objects in Tigris?\n\nTurns out Git packfiles and how they intersected with my (admittedly somewhat\nterrible) filesystem shim were the main reason why it was slow. I ended up\nhaving to invent my own packfile format with a columnar store that’s object\nstorage native. This is the fruit of all of my performance analysis, metrics\nannotations, and more\n[Texas-style distributed systems work](https://www.usenix.org/system/files/1311_05-08_mickens.pdf)\nthan you can make your k8s cluster shake sticks at.\n\nThis approach worked surprisingly well for production-sized repositories, so I’m sticking with this new Packfile format for now. It seems the least obtrusive change to make Git Objects feel like object storage Objects, without any client side changes.\n\n## What is a Git? [A miserable little pile of objects!](https://youtu.be/5tV33Ewf_hw)[](#what-is-a-git-a-miserable-little-pile-of-objects)\n\nWhen you make a commit, Git stores the changes you make as objects inside the\n`.git` (I’ll call this “dotgit” so I don’t have to write as many backticks)\nfolder.\n\nImagine Git as two things: a sea of objects and named references to individual\nobjects. Each object is a\n[content-addressed](https://en.wikipedia.org/wiki/Content-addressable_storage)\nand compressed file. Here’s an example from a tiny git repository:\n\n```\n$ mkdir ~/tmp/gitexample\n$ git init && git branch -m main\n$ echo \"Hello, blog!\" >> hello.txt\n$ git add .\n$ git commit -sm \"chore: initial commit\"\n```\nThis produces several objects on the disk like this:\n\nIf you want to read the contents of an object, it’s compressed, so you have to use a fairly evil looking python oneliner to scoop out the tasty innards:\n\n```\n$ file .git/objects/**/* | grep -v directory\n.git/objects/1c/7a26a901724b4ce766655ac387413fb9ec7966: zlib compressed data\n.git/objects/8e/67afbb2ee6bdcbb79061dfdfb93febce857bd3: zlib compressed data\n.git/objects/9c/c9867337c2ebae85ba2350f901e0bcc209fe26: zlib compressed data\n$ python3 -c \"import sys, zlib; sys.stdout.buffer.write(zlib.decompress(sys.stdin.buffer.read()))\" < .git/objects/9c/c9867337c2ebae85ba2350f901e0bcc209fe26\nblob 13Hello, blog!\n```\nAs you can see, the objects are just bare files. Let’s look at a Git repository of the Linux kernel and try to extract out an arbitrary commit. Everything should just be a billionty bare object files, right? It should be easy to find a single commit just by looking for the ID on the disk, right?\n\nIf only reality were so simple:\n\n```\n$ cd ~/Code/linux.git/\n$ tree objects\nobjects\n├── info\n└── pack\n    ├── pack-45986f41063f286029742ec12e2c2882b88c5786.idx\n    ├── pack-45986f41063f286029742ec12e2c2882b88c5786.pack\n    └── pack-45986f41063f286029742ec12e2c2882b88c5786.rev\n3 directories, 3 files\n```\nYeah, as I’m sure you guessed just putting everything into their own files won’t\nscale to something like the Linux kernel. I’m pretty sure you’d run into inode\nlimits like everyone did in the era of\n[fractal `node_modules` folders](https://github.com/npm/npm/issues/11747).\n\nIf you’ve used Node for long enough to remember that, please go get a colonoscopy. Colon cancer is a real concern that too many people overlook for too long and takes too many lives too early.\n\nGit works around this by putting objects into\n[packfiles](https://git-scm.com/book/en/v2/Git-Internals-Packfiles), compressed\nbundles of objects that store them all in the same file. Here's an example of\nthe packfile efficiency in my checkout of objgit:\n\n```\n$ git count-objects -v\ncount: 756\nsize: 3500\nin-pack: 448\npacks: 1\nsize-pack: 321\nprune-packable: 0\ngarbage: 0\nsize-garbage: 0\n```\nIf you ever need to “force” git to put bare objects into a packfile, you can run\n`git gc`:\n\n```\n$ git gc\n[omitted for brevity]\n$ git count-objects -v\ncount: 0\nsize: 0\nin-pack: 1203\npacks: 2\nsize-pack: 848\nprune-packable: 0\ngarbage: 0\nsize-garbage: 0\n```\nOne of the beautiful things about implementing Git on top of object storage like I am is that I’m using a platform where the object data is a sea of objects with named references to points in that sea stored in FoundationDB. This is a kind of divine recursion that I don’t really know how to describe the beauty of. As above, so below.\n\n## Yo dawg, herd you like objects[](#yo-dawg-herd-you-like-objects)\n\nHere's the object count for a copy of the Linux kernel:\n\n```\nxe@zohar:~/Code/linux.git$ git count-objects -v\ncount: 0\nsize: 0\nin-pack: 11827138\npacks: 1\nsize-pack: 3876775\nprune-packable: 0\ngarbage: 0\nsize-garbage: 0\n```\nThis is eleven million objects, which at a very generous assumption of 10ms per GetObject call means that fetching each of them takes over an hour to fetch them all. The truth is there really aren’t 11M objects as individual files on the disk, they’re bundled into one big happy 3.4Gi packfile. Your typical git repo ends up accumulating them as it makes sense to break them up. My local copy of the Tigris blog has 4 packfiles and 290-ish bare objects.\n\n**So you’d be thinking, “Oh, if git has packfiles, then why is the rest of this\npost a thing?”**\n\nWell, like many things in distributed systems it’s complicated. Packfiles are\ndifficult because they’re designed with local storage and/or mmap in mind. Git\nconstantly writes packfiles to disk and then re-reads them. Filesystem reads in\nthat case are 10 nanoseconds *at most* (the filesystem cache helps so much here)\nbut doing any network roundtrip is 10 milliseconds *at minimum*. It’s at least a\nmillion times slower because of how reality works.\n\nOne of the things that `/usr/bin/git` does that makes integrating it into object\nstorage difficult is the unix-y idiom of writing to a file and then immediately\nreading back from that file to calculate the hash. In object storage you can’t\nGetObject something that hasn’t finished a PutObject call. I worked around this\npreviously by writing to the disk and then doing it that way, but the experience\nkinda sucked in practice.\n\n### Messin' with Packfiles[](#messin-with-packfiles)\n\nEach packfile has an index that describes what’s in it. Here’s a view of the index of the packfile made out of that trivial Git repo from earlier uppost:\n\n```\n$ git gc # force objects into a packfile\n$ git verify-pack -v .git/objects/pack/pack-3971f5085c23c38be00e517ed0c64ca7df19b746.idx\n1c7a26a901724b4ce766655ac387413fb9ec7966 commit 526 366 12\n9cc9867337c2ebae85ba2350f901e0bcc209fe26 blob   13 22 378\n8e67afbb2ee6bdcbb79061dfdfb93febce857bd3 tree   37 48 400\nnon delta: 3 objects\n.git/objects/pack/pack-3971f5085c23c38be00e517ed0c64ca7df19b746.pack: ok\n```\nThe commit points to the tree whose file “hello.txt” points to the blob and, blob’s your uncle, you have a repo. Git uses these binary indices to let it know where to look and how far it needs to seek into the packfile to know where to go to get things.\n\nThe great part is that this works really well when everything is in a\nfilesystem. Git [mmaps](https://en.wikipedia.org/wiki/Mmap) the packfiles so\nthat the kernel treats disk contents as memory pages, meaning that trying to\nread past what’s “in memory” makes the kernel load it instead of userspace. This\nis faster than loading it from the disk directly. It’s a shame this design\ndoesn’t work in object storage.\n\n### If only you could construct Range requests from packfiles[](#if-only-you-could-construct-range-requests-from-packfiles)\n\nAt some level this sounds pretty great for object storage, right? You have\noffsets into the packfiles and then you can “just” grab out a single object from\na packfile with an\n[HTTP Range request](https://developer.mozilla.org/en-US/docs/Web/HTTP/Guides/Range_requests)\nright? Objects are placed randomly within packfiles and other attempts at\nstoring Git in object storage end up having problems here. Tigris is really good\nat random access scans, so most of the hard part is figuring out how to grab the\nright data out of the bucket.\n\nHTTP Range requests let a client download *part* of a file. The main usecase\nthey're built for is back in the day of dial-up internet you weren't online all\nthe time. Your main path to the Internet was the same way you send and received\nphone calls. As such, if someone called you while you were online, all your\ndownloads got interrupted. Range requests let Internet Explorer resume downloads\nwhere they got cut off instead of having to start all the way over.\n\nThey've been maintained into the modern era but don't really get much use outside of galaxy brain format abuse like what I'm doing and online video streaming.\n\nWell, it’s complicated. The example I gave shows all of the index entries one\nafter the other, but in the real world processing the index entries one after\nthe other you know the *decompressed* size of a single object in the packfile,\nbut not the *compressed* size. This means you don’t have enough information to\nconstruct a HTTP Range request.\n\nThis core problem is half the reason why I ended up needing to make my own object-storage native Git packfile format.\n\n## SEND CUE SHEET[](#send-cue-sheet)\n\nWay back in the days of physical media, one of the most common formats was the CD-ROM (Compact Disc Read-Only-Memory, or CD). A CD is a 700Mi container that stores data in sessions that each contain up to 99 tracks of either audio or data. CDs were originally invented to store song audio in so that you could listen to an hour or so of music at a higher quality than analogue cassette tapes. CDs also let the player skip from track to track so you can go directly to the song or movement of a larger work that you like.\n\nOf course this backfires if your CD mastering team decided to put\n[the entirety of Dancing Mad](https://youtu.be/g6UgxJtHIVY?list=RDg6UgxJtHIVY)\ninto a single 17 minute track, meaning that you just have to know that\n[the blessed fourth movement](https://youtu.be/g6UgxJtHIVY?t=590) is about 9\nminutes into the battle music.\n\nThis is also why you see guides telling you to put legal backups of CD and DVD\nmedia into lossless `.iso` files. An `.iso` file contains one recording session\nthat may contain data or audio.\n\nHowever there’s one catch that kinda ruins this easy way to back up CDs: they\ncan store *multiple* recording sessions on the same disc. Most of the time this\nwasn’t used outside of making piracy on certain late 90’s/early 00’s game\nconsoles more annoying, but there was\n[that one Ricoh Encryptease product](https://youtu.be/_5ucImqdKbY) that combined\na user-recordable area with a factory printed area so that you could encrypt\nfiles on CDs you share with the decryption software shipping alongside it. This\nis about as cursed as it sounds.\n\nThe trick of using multiple sessions is how Dreamcast games play as audio CDs telling you to put it into a Dreamcast or how Xbox 360 games play as DVDs telling you to put it into an Xbox 360. As an added bonus it means that when you stick it into a computer it thinks that it’s an audio CD or DVD, which means that lazy pirates can’t easily scoop out all the game files to their hard drives.\n\nAs a result, there needed to be a way to properly handle this for archival\npurposes. The eventual result was creating\n[cue sheets](<https://en.wikipedia.org/wiki/Cue_sheet_(computing)>) to store\nalongside the binary blob of data. The `.cue` sheet stores information that the\ndecoder uses to be able to seek to arbitrary points in the `.bin` file. This\nlets you easily extract things like songs or bits of data without having to read\nthe entire CD image. As an added bonus it handles multiple recording sessions\nfor you.\n\n## Packfiles v2: object storage boogaloo[](#packfiles-v2-object-storage-boogaloo)\n\nThis got me thinking, how would we take all of these lessons into heart and build a new git packfile format optimized for object storage?\n\nWait, I know what you’re thinking. You’re thinking that I’m about to make a\n[Chesterton’s Fence violation](https://en.wikipedia.org/wiki/Wikipedia:Chesterton%27s_fence).\nJust “rolling my own” format for something as dear and precious as storing the\nrevision history of a company’s code repositories is probably one of the worst\ndecisions you can make, right?\n\nNormally, yes, it’s a bad idea to do this. However Git is a *distributed*\nversion control system. When you clone a repository, you clone *all* of the\nchanges ever made to it on every branch at the same time. This also means that\neveryone has a copy of the entire history of that repository, meaning that if\nthe worst does in fact come to pass and my handrolled format ends up sucking\nit’s trivial to recreate all the data. Just push it again.\n\nSo what would this format look like?\n\nWell for one the format needs to be Range-request native. You should be able to\nscoop any one object out of a packfile without having to download or process\nanything but the object you want. Again, Tigris is good at this, so we should\ndesign the format with that usecase directly in mind. The format should also\ntake advantage of modern compression libraries like\n[zstd](https://en.wikipedia.org/wiki/Zstd) which are faster and more\ndata-efficient than zlib. Finally delta objects should be stored as their own\nobject in the packfile instead of slapped onto the end of the object it’s a\ndelta of so that you don’t have to read the object and its deltas to read the\nobject in the first place.\n\nI skipped over this earlier to save time, but Git stores both file revisions (the entire copy of a file at any given point in time) and the difference between them as an optimization to make it easier to uncompute the changes made in commits. At some level this meme is both accurate and wrong:\n\n### Objgit's packfile format that probably needs a name[](#objgits-packfile-format-that-probably-needs-a-name)\n\nThe format I came up with is pretty directly inspired from the `.bin` and `.cue`\nformat of CD backup. Objects are stored one after the other in a `.bin` file\nthat’s normally up to 128Mi (the oddly specific number was chosen because it\nlooked round to me) and the metadata of what objects are in there are stored\nseparately in a binary-encoded `.cue` sheet. Together this makes Git repository\nstorage in Tigris a columnar store.\n\nThe big thing I did was store the sizes for both the *compressed* and\n*uncompressed* forms of objects alongside the offset into the packfile. This\nmeans that you can trivially construct the right HTTP Range requests to scoop\nindividual objects out of Tigris while the packfile is downloading in the\nbackground.\n\nAs an optimization for latency, whenever the Git library requests any object from a packfile, the entire packfile is downloaded from object storage to a temporary folder. Anything on the “far end” of the packfile is Range-requested from Tigris until the downloaded packfile “catches up”. In practice this ends up meaning that packfiles get downloaded just in time for them to be useful to read from and there’s overall fairly little latency beyond what’s unavoidable with the Git library I’m using. If the “scooping objects out of the far end” problem ends up being an issue in practice I’ll just have it start fetching the most recent packfiles for a given repository in the background when you start pushing or pulling.\n\nThere’s nothing in the definition of the packfile format that limits packfiles\nto 128Mi, I’m just doing that to prevent them from getting too big to download\nquickly. In theory if you store a large binary blob (such as 3d models,\nperfectly legal backups, etc) into the git repository directly it could result\nin a packfile that’s bigger than 128Mi. I plan to solve this by implementing\n[Git Large File Storage](https://git-lfs.com/) in the near future.\n\nBut for now if you have a workflow that involves storing large binary files in your git repository and want to use objgit for that: consider a different architecture.\n\n## The funny numbers[](#the-funny-numbers)\n\nOne of the more surprising things about Git is that basically every interaction with it is expensive. As such, you can go a long way by benchmarking how long it takes to push/pull repositories. As such, I decided to compare against a few git repos that have some interesting properties:\n\n- objgit itself: [https://github.com/tigrisdata/objgit](https://github.com/tigrisdata/objgit) , a newer repository with\nsome AI assisted development for prototyping.\n- My e/x/perimental monorepo: [https://github.com/Xe/x/](https://github.com/Xe/x/) , a repo that's been\naround for over a decade and holds many experiments.\n- The Tigris blog repo: [https://github.com/tigrisdata/tigris-blog](https://github.com/tigrisdata/tigris-blog) , a repo that\nhas a lot of images and other incompressible objects in it.\n\nHere are the conditions I ran the tests in:\n\n| Item | Value | \n|---|---|\n| Started | 2026-09-11T12:28:59-04:00 | \n| Finished | 2026-09-11T13:06:52-04:00 | \n| Host | Mac, 16 CPUs | \n| Go | go1.26.5 | \n| Git | git version 2.55.0 | \n| Bucket | xe-objgit-devel | \n| Build \"old\" | v1.0.2 | \n| Build \"fixed\" | origin/main | \n\n### Push tests[](#push-tests)\n\nThe big thing I wanted to fix was reducing the number of object storage calls. Less object storage calls, less latency waiting for them to resolve. It ended up being ridiculously effective. Object storage call counts sank like a stone.\n\nAll charts in this post are at logarithmic scales so they render more cleanly and the green bars are visible.\n\n| Repo | Build | Wall | S3 requests | PUT | GET | HEAD | LIST | Keys | Bucket bytes | \n|---|---|---|---|---|---|---|---|---|---|\n| objgit | old | 8.7s (8.7s-15.9s) | 231 | 46 | 47 | 0 | 138 | 42 | 829.27 KiB | \n| objgit | new | 2.2s (1.9s-2.5s) | 18 (17-20) | 5 | 10 (9-12) | 0 | 3 | 4 | 902.98 KiB | \n| x | old | 3m29.4s (3m29.4s-3m40.6s) | 9,236 (9,236-9,304) | 1,087 | 2,170 | 0 | 5,979 (5,979-6,047) | 1,082 | 54.96 MiB | \n| x | new | 14.3s (14.3s-14.7s) | 30 | 6 | 20 | 0 | 4 | 4 | 46.38 MiB | \n| tigris-blog | old | 2m13.4s (1m29.1s-2m30.5s) | 3,324 (2,687-3,417) | 515 | 522 | 0 | 2,287 (1,650-2,380) | 511 | 354.67 MiB | \n| tigris-blog | new | 26.5s (19.2s-27.4s) | 136 (113-146) | 9 | 123 (100-133) | 0 | 4 | 8 | 360.75 MiB | \n\nThe biggest gain was wall clock time for pushing though:\n\nOne of the biggest places that objgit used to lag was pushing taking way longer than it felt like it should. Eliminating the Tigris round trips made pushing way more responsive.\n\n### Clone tests[](#clone-tests)\n\nI also wanted to see how the difference affected clone times. There were the same benefits as with pushing:\n\n| Repo | Build | Wall | S3 requests | GET | HEAD | LIST | Wire bytes | \n|---|---|---|---|---|---|---|---|\n| objgit | old | 11.8s (11.5s-19.5s) | 323 | 51 | 0 | 272 | 763.63 KiB | \n| objgit | new | 2.6s (1.5s-5.7s) | 17 (16-17) | 13 (12-13) | 0 | 4 | 767.96 KiB | \n| x | old | 3m23.5s (3m23.5s-3m33.6s) | 6,428 (6,428-6,780) | 1,091 | 0 | 5,337 (5,337-5,689) | 40.68 MiB (40.68 MiB-40.70 MiB) | \n| x | new | 54.4s (54.4s-57.8s) | 17 (17-71) | 13 (13-67) | 0 | 4 | 42.22 MiB (42.22 MiB-42.36 MiB) | \n| tigris-blog | old | 2m23.6s (2m19.4s-2m50.2s) | 3,675 (3,601-4,123) | 520 | 0 | 3,155 (3,081-3,603) | 350.94 MiB (350.93 MiB-351.07 MiB) | \n| tigris-blog | new | 1m22s (1m20.8s-1m38.7s) | 158 (137-317) | 155 (134-314) | 0 | 3 | 349.04 MiB (348.00 MiB-352.67 MiB) | \n\nOh yeah, the time numbers would probably be better if I tested this on a machine with ethernet. I did all my testing with my corp laptop on Wi-Fi to specifically put this in one of the worst conditions it could possibly be in.\n\n## Conclusion section[](#conclusion-section)\n\nI’m still actively working on this. I’m not confident enough to use this for my\nown projects yet and I wouldn’t blame you for not wanting to use this yet\neither. I still haven’t implemented authentication, authorization, any kind of\nAPI (my long-form [SigV4 auth post](https://www.tigrisdata.com/blog/sigv4/) was\nactually going to be an objgit post!), or any rate limit beyond what your\nmachine can physically process. If you were to take this, run it, and then\nexpose it to the Internet, then anyone that can connect to that server can pull\nor push whatever they want. Consider not doing that.\n\nObjgit packfiles also currently accumulate forever, so if you have a bunch of small pushes then there will be a bunch of small packfiles in the bucket. I’m toying with designs that would occasionally compact them into bigger packfiles, but that’s something that can be done later.\n\nAt the least though: Git’s packfile format is a great format for the constraints of storing git repositories in actual filesystems. The moment you put network roundtrips into the mix it all goes south.\n\nI'm gonna keep working on this and publish reports like this as I learn more. I hope this was interesting! Stay safe out there.\n\nobjgit stores git repositories directly in Tigris. The .bin/.cue container, the columnar index, and the four-tier read ladder in this post are all in the repo.","body_html":"<h1 id=\"you-can-run-git-on-object-storage-if-you-re-make-packfiles\">You can run git on object storage if you re-make packfiles</h1>\n<p><a href=\"/blog/\">Blog</a></p>\n<h1 id=\"you-can-run-git-on-object-storage-if-you-re-make-packfiles-2\">You can run git on object storage if you re-make packfiles</h1>\n<p>Git packfiles were designed for mmap and local disk, so pulling one object out of a bucket means guessing at a byte range. I wrote a new format instead.</p>\n<h2 id=\"contents\">Contents</h2>\n<p>It sure seems that a bunch of companies are trying to ship a git product of some kind as of late. Wonder why that is.</p>\n<p>Either way, I’m building a Git server backed by object storage as an\n<a href=\"https://github.com/tigrisdata/objgit\" rel=\"nofollow ugc noopener\">open-source project</a>. It sounded simple\nenough to start: Git looks like a filesystem, so let’s use a filesystem as a\ntranslation layer on top of object storage to make Git speak object storage.\nThis model worked… <a href=\"https://www.tigrisdata.com/blog/objgit/\" rel=\"nofollow ugc noopener\">ok, I guess?</a> But\nit didn’t work for real-world size repositories, so I needed a different\napproach. Git stores everything in\n<a href=\"https://git-scm.com/book/en/v2/Git-Internals-Git-Objects\" rel=\"nofollow ugc noopener\">Objects</a>, so why not\nstore those as objects in Tigris?</p>\n<p>Turns out Git packfiles and how they intersected with my (admittedly somewhat\nterrible) filesystem shim were the main reason why it was slow. I ended up\nhaving to invent my own packfile format with a columnar store that’s object\nstorage native. This is the fruit of all of my performance analysis, metrics\nannotations, and more\n<a href=\"https://www.usenix.org/system/files/1311_05-08_mickens.pdf\" rel=\"nofollow ugc noopener\">Texas-style distributed systems work</a>\nthan you can make your k8s cluster shake sticks at.</p>\n<p>This approach worked surprisingly well for production-sized repositories, so I’m sticking with this new Packfile format for now. It seems the least obtrusive change to make Git Objects feel like object storage Objects, without any client side changes.</p>\n<h2 id=\"what-is-a-git-a-miserable-little-pile-of-objects\">What is a Git? <a href=\"https://youtu.be/5tV33Ewf_hw\" rel=\"nofollow ugc noopener\">A miserable little pile of objects!</a><a href=\"#what-is-a-git-a-miserable-little-pile-of-objects\"></a></h2>\n<p>When you make a commit, Git stores the changes you make as objects inside the\n<code>.git</code> (I’ll call this “dotgit” so I don’t have to write as many backticks)\nfolder.</p>\n<p>Imagine Git as two things: a sea of objects and named references to individual\nobjects. Each object is a\n<a href=\"https://en.wikipedia.org/wiki/Content-addressable_storage\" rel=\"nofollow ugc noopener\">content-addressed</a>\nand compressed file. Here’s an example from a tiny git repository:</p>\n<pre><code>$ mkdir ~/tmp/gitexample\n$ git init &amp;&amp; git branch -m main\n$ echo &quot;Hello, blog!&quot; &gt;&gt; hello.txt\n$ git add .\n$ git commit -sm &quot;chore: initial commit&quot;</code></pre>\n<p>This produces several objects on the disk like this:</p>\n<p>If you want to read the contents of an object, it’s compressed, so you have to use a fairly evil looking python oneliner to scoop out the tasty innards:</p>\n<pre><code>$ file .git/objects/**/* | grep -v directory\n.git/objects/1c/7a26a901724b4ce766655ac387413fb9ec7966: zlib compressed data\n.git/objects/8e/67afbb2ee6bdcbb79061dfdfb93febce857bd3: zlib compressed data\n.git/objects/9c/c9867337c2ebae85ba2350f901e0bcc209fe26: zlib compressed data\n$ python3 -c &quot;import sys, zlib; sys.stdout.buffer.write(zlib.decompress(sys.stdin.buffer.read()))&quot; &lt; .git/objects/9c/c9867337c2ebae85ba2350f901e0bcc209fe26\nblob 13Hello, blog!</code></pre>\n<p>As you can see, the objects are just bare files. Let’s look at a Git repository of the Linux kernel and try to extract out an arbitrary commit. Everything should just be a billionty bare object files, right? It should be easy to find a single commit just by looking for the ID on the disk, right?</p>\n<p>If only reality were so simple:</p>\n<pre><code>$ cd ~/Code/linux.git/\n$ tree objects\nobjects\n├── info\n└── pack\n    ├── pack-45986f41063f286029742ec12e2c2882b88c5786.idx\n    ├── pack-45986f41063f286029742ec12e2c2882b88c5786.pack\n    └── pack-45986f41063f286029742ec12e2c2882b88c5786.rev\n3 directories, 3 files</code></pre>\n<p>Yeah, as I’m sure you guessed just putting everything into their own files won’t\nscale to something like the Linux kernel. I’m pretty sure you’d run into inode\nlimits like everyone did in the era of\n<a href=\"https://github.com/npm/npm/issues/11747\" rel=\"nofollow ugc noopener\">fractal <code>node_modules</code> folders</a>.</p>\n<p>If you’ve used Node for long enough to remember that, please go get a colonoscopy. Colon cancer is a real concern that too many people overlook for too long and takes too many lives too early.</p>\n<p>Git works around this by putting objects into\n<a href=\"https://git-scm.com/book/en/v2/Git-Internals-Packfiles\" rel=\"nofollow ugc noopener\">packfiles</a>, compressed\nbundles of objects that store them all in the same file. Here&#39;s an example of\nthe packfile efficiency in my checkout of objgit:</p>\n<pre><code>$ git count-objects -v\ncount: 756\nsize: 3500\nin-pack: 448\npacks: 1\nsize-pack: 321\nprune-packable: 0\ngarbage: 0\nsize-garbage: 0</code></pre>\n<p>If you ever need to “force” git to put bare objects into a packfile, you can run\n<code>git gc</code>:</p>\n<pre><code>$ git gc\n[omitted for brevity]\n$ git count-objects -v\ncount: 0\nsize: 0\nin-pack: 1203\npacks: 2\nsize-pack: 848\nprune-packable: 0\ngarbage: 0\nsize-garbage: 0</code></pre>\n<p>One of the beautiful things about implementing Git on top of object storage like I am is that I’m using a platform where the object data is a sea of objects with named references to points in that sea stored in FoundationDB. This is a kind of divine recursion that I don’t really know how to describe the beauty of. As above, so below.</p>\n<h2 id=\"yo-dawg-herd-you-like-objects\">Yo dawg, herd you like objects<a href=\"#yo-dawg-herd-you-like-objects\"></a></h2>\n<p>Here&#39;s the object count for a copy of the Linux kernel:</p>\n<pre><code>xe@zohar:~/Code/linux.git$ git count-objects -v\ncount: 0\nsize: 0\nin-pack: 11827138\npacks: 1\nsize-pack: 3876775\nprune-packable: 0\ngarbage: 0\nsize-garbage: 0</code></pre>\n<p>This is eleven million objects, which at a very generous assumption of 10ms per GetObject call means that fetching each of them takes over an hour to fetch them all. The truth is there really aren’t 11M objects as individual files on the disk, they’re bundled into one big happy 3.4Gi packfile. Your typical git repo ends up accumulating them as it makes sense to break them up. My local copy of the Tigris blog has 4 packfiles and 290-ish bare objects.</p>\n<p><strong>So you’d be thinking, “Oh, if git has packfiles, then why is the rest of this\npost a thing?”</strong></p>\n<p>Well, like many things in distributed systems it’s complicated. Packfiles are\ndifficult because they’re designed with local storage and/or mmap in mind. Git\nconstantly writes packfiles to disk and then re-reads them. Filesystem reads in\nthat case are 10 nanoseconds <em>at most</em> (the filesystem cache helps so much here)\nbut doing any network roundtrip is 10 milliseconds <em>at minimum</em>. It’s at least a\nmillion times slower because of how reality works.</p>\n<p>One of the things that <code>/usr/bin/git</code> does that makes integrating it into object\nstorage difficult is the unix-y idiom of writing to a file and then immediately\nreading back from that file to calculate the hash. In object storage you can’t\nGetObject something that hasn’t finished a PutObject call. I worked around this\npreviously by writing to the disk and then doing it that way, but the experience\nkinda sucked in practice.</p>\n<h3 id=\"messin-with-packfiles\">Messin&#39; with Packfiles<a href=\"#messin-with-packfiles\"></a></h3>\n<p>Each packfile has an index that describes what’s in it. Here’s a view of the index of the packfile made out of that trivial Git repo from earlier uppost:</p>\n<pre><code>$ git gc # force objects into a packfile\n$ git verify-pack -v .git/objects/pack/pack-3971f5085c23c38be00e517ed0c64ca7df19b746.idx\n1c7a26a901724b4ce766655ac387413fb9ec7966 commit 526 366 12\n9cc9867337c2ebae85ba2350f901e0bcc209fe26 blob   13 22 378\n8e67afbb2ee6bdcbb79061dfdfb93febce857bd3 tree   37 48 400\nnon delta: 3 objects\n.git/objects/pack/pack-3971f5085c23c38be00e517ed0c64ca7df19b746.pack: ok</code></pre>\n<p>The commit points to the tree whose file “hello.txt” points to the blob and, blob’s your uncle, you have a repo. Git uses these binary indices to let it know where to look and how far it needs to seek into the packfile to know where to go to get things.</p>\n<p>The great part is that this works really well when everything is in a\nfilesystem. Git <a href=\"https://en.wikipedia.org/wiki/Mmap\" rel=\"nofollow ugc noopener\">mmaps</a> the packfiles so\nthat the kernel treats disk contents as memory pages, meaning that trying to\nread past what’s “in memory” makes the kernel load it instead of userspace. This\nis faster than loading it from the disk directly. It’s a shame this design\ndoesn’t work in object storage.</p>\n<h3 id=\"if-only-you-could-construct-range-requests-from-packfiles\">If only you could construct Range requests from packfiles<a href=\"#if-only-you-could-construct-range-requests-from-packfiles\"></a></h3>\n<p>At some level this sounds pretty great for object storage, right? You have\noffsets into the packfiles and then you can “just” grab out a single object from\na packfile with an\n<a href=\"https://developer.mozilla.org/en-US/docs/Web/HTTP/Guides/Range_requests\" rel=\"nofollow ugc noopener\">HTTP Range request</a>\nright? Objects are placed randomly within packfiles and other attempts at\nstoring Git in object storage end up having problems here. Tigris is really good\nat random access scans, so most of the hard part is figuring out how to grab the\nright data out of the bucket.</p>\n<p>HTTP Range requests let a client download <em>part</em> of a file. The main usecase\nthey&#39;re built for is back in the day of dial-up internet you weren&#39;t online all\nthe time. Your main path to the Internet was the same way you send and received\nphone calls. As such, if someone called you while you were online, all your\ndownloads got interrupted. Range requests let Internet Explorer resume downloads\nwhere they got cut off instead of having to start all the way over.</p>\n<p>They&#39;ve been maintained into the modern era but don&#39;t really get much use outside of galaxy brain format abuse like what I&#39;m doing and online video streaming.</p>\n<p>Well, it’s complicated. The example I gave shows all of the index entries one\nafter the other, but in the real world processing the index entries one after\nthe other you know the <em>decompressed</em> size of a single object in the packfile,\nbut not the <em>compressed</em> size. This means you don’t have enough information to\nconstruct a HTTP Range request.</p>\n<p>This core problem is half the reason why I ended up needing to make my own object-storage native Git packfile format.</p>\n<h2 id=\"send-cue-sheet\">SEND CUE SHEET<a href=\"#send-cue-sheet\"></a></h2>\n<p>Way back in the days of physical media, one of the most common formats was the CD-ROM (Compact Disc Read-Only-Memory, or CD). A CD is a 700Mi container that stores data in sessions that each contain up to 99 tracks of either audio or data. CDs were originally invented to store song audio in so that you could listen to an hour or so of music at a higher quality than analogue cassette tapes. CDs also let the player skip from track to track so you can go directly to the song or movement of a larger work that you like.</p>\n<p>Of course this backfires if your CD mastering team decided to put\n<a href=\"https://youtu.be/g6UgxJtHIVY?list=RDg6UgxJtHIVY\" rel=\"nofollow ugc noopener\">the entirety of Dancing Mad</a>\ninto a single 17 minute track, meaning that you just have to know that\n<a href=\"https://youtu.be/g6UgxJtHIVY?t=590\" rel=\"nofollow ugc noopener\">the blessed fourth movement</a> is about 9\nminutes into the battle music.</p>\n<p>This is also why you see guides telling you to put legal backups of CD and DVD\nmedia into lossless <code>.iso</code> files. An <code>.iso</code> file contains one recording session\nthat may contain data or audio.</p>\n<p>However there’s one catch that kinda ruins this easy way to back up CDs: they\ncan store <em>multiple</em> recording sessions on the same disc. Most of the time this\nwasn’t used outside of making piracy on certain late 90’s/early 00’s game\nconsoles more annoying, but there was\n<a href=\"https://youtu.be/_5ucImqdKbY\" rel=\"nofollow ugc noopener\">that one Ricoh Encryptease product</a> that combined\na user-recordable area with a factory printed area so that you could encrypt\nfiles on CDs you share with the decryption software shipping alongside it. This\nis about as cursed as it sounds.</p>\n<p>The trick of using multiple sessions is how Dreamcast games play as audio CDs telling you to put it into a Dreamcast or how Xbox 360 games play as DVDs telling you to put it into an Xbox 360. As an added bonus it means that when you stick it into a computer it thinks that it’s an audio CD or DVD, which means that lazy pirates can’t easily scoop out all the game files to their hard drives.</p>\n<p>As a result, there needed to be a way to properly handle this for archival\npurposes. The eventual result was creating\n<a href=\"https://en.wikipedia.org/wiki/Cue_sheet_(computing)\" rel=\"nofollow ugc noopener\">cue sheets</a> to store\nalongside the binary blob of data. The <code>.cue</code> sheet stores information that the\ndecoder uses to be able to seek to arbitrary points in the <code>.bin</code> file. This\nlets you easily extract things like songs or bits of data without having to read\nthe entire CD image. As an added bonus it handles multiple recording sessions\nfor you.</p>\n<h2 id=\"packfiles-v2-object-storage-boogaloo\">Packfiles v2: object storage boogaloo<a href=\"#packfiles-v2-object-storage-boogaloo\"></a></h2>\n<p>This got me thinking, how would we take all of these lessons into heart and build a new git packfile format optimized for object storage?</p>\n<p>Wait, I know what you’re thinking. You’re thinking that I’m about to make a\n<a href=\"https://en.wikipedia.org/wiki/Wikipedia:Chesterton%27s_fence\" rel=\"nofollow ugc noopener\">Chesterton’s Fence violation</a>.\nJust “rolling my own” format for something as dear and precious as storing the\nrevision history of a company’s code repositories is probably one of the worst\ndecisions you can make, right?</p>\n<p>Normally, yes, it’s a bad idea to do this. However Git is a <em>distributed</em>\nversion control system. When you clone a repository, you clone <em>all</em> of the\nchanges ever made to it on every branch at the same time. This also means that\neveryone has a copy of the entire history of that repository, meaning that if\nthe worst does in fact come to pass and my handrolled format ends up sucking\nit’s trivial to recreate all the data. Just push it again.</p>\n<p>So what would this format look like?</p>\n<p>Well for one the format needs to be Range-request native. You should be able to\nscoop any one object out of a packfile without having to download or process\nanything but the object you want. Again, Tigris is good at this, so we should\ndesign the format with that usecase directly in mind. The format should also\ntake advantage of modern compression libraries like\n<a href=\"https://en.wikipedia.org/wiki/Zstd\" rel=\"nofollow ugc noopener\">zstd</a> which are faster and more\ndata-efficient than zlib. Finally delta objects should be stored as their own\nobject in the packfile instead of slapped onto the end of the object it’s a\ndelta of so that you don’t have to read the object and its deltas to read the\nobject in the first place.</p>\n<p>I skipped over this earlier to save time, but Git stores both file revisions (the entire copy of a file at any given point in time) and the difference between them as an optimization to make it easier to uncompute the changes made in commits. At some level this meme is both accurate and wrong:</p>\n<h3 id=\"objgit-s-packfile-format-that-probably-needs-a-name\">Objgit&#39;s packfile format that probably needs a name<a href=\"#objgits-packfile-format-that-probably-needs-a-name\"></a></h3>\n<p>The format I came up with is pretty directly inspired from the <code>.bin</code> and <code>.cue</code>\nformat of CD backup. Objects are stored one after the other in a <code>.bin</code> file\nthat’s normally up to 128Mi (the oddly specific number was chosen because it\nlooked round to me) and the metadata of what objects are in there are stored\nseparately in a binary-encoded <code>.cue</code> sheet. Together this makes Git repository\nstorage in Tigris a columnar store.</p>\n<p>The big thing I did was store the sizes for both the <em>compressed</em> and\n<em>uncompressed</em> forms of objects alongside the offset into the packfile. This\nmeans that you can trivially construct the right HTTP Range requests to scoop\nindividual objects out of Tigris while the packfile is downloading in the\nbackground.</p>\n<p>As an optimization for latency, whenever the Git library requests any object from a packfile, the entire packfile is downloaded from object storage to a temporary folder. Anything on the “far end” of the packfile is Range-requested from Tigris until the downloaded packfile “catches up”. In practice this ends up meaning that packfiles get downloaded just in time for them to be useful to read from and there’s overall fairly little latency beyond what’s unavoidable with the Git library I’m using. If the “scooping objects out of the far end” problem ends up being an issue in practice I’ll just have it start fetching the most recent packfiles for a given repository in the background when you start pushing or pulling.</p>\n<p>There’s nothing in the definition of the packfile format that limits packfiles\nto 128Mi, I’m just doing that to prevent them from getting too big to download\nquickly. In theory if you store a large binary blob (such as 3d models,\nperfectly legal backups, etc) into the git repository directly it could result\nin a packfile that’s bigger than 128Mi. I plan to solve this by implementing\n<a href=\"https://git-lfs.com/\" rel=\"nofollow ugc noopener\">Git Large File Storage</a> in the near future.</p>\n<p>But for now if you have a workflow that involves storing large binary files in your git repository and want to use objgit for that: consider a different architecture.</p>\n<h2 id=\"the-funny-numbers\">The funny numbers<a href=\"#the-funny-numbers\"></a></h2>\n<p>One of the more surprising things about Git is that basically every interaction with it is expensive. As such, you can go a long way by benchmarking how long it takes to push/pull repositories. As such, I decided to compare against a few git repos that have some interesting properties:</p>\n<ul><li><p>objgit itself: <a href=\"https://github.com/tigrisdata/objgit\" rel=\"nofollow ugc noopener\"><a href=\"https://github.com/tigrisdata/objgit\" rel=\"nofollow ugc noopener\">https://github.com/tigrisdata/objgit</a></a> , a newer repository with</p><p>some AI assisted development for prototyping.</p></li><li><p>My e/x/perimental monorepo: <a href=\"https://github.com/Xe/x/\" rel=\"nofollow ugc noopener\"><a href=\"https://github.com/Xe/x/\" rel=\"nofollow ugc noopener\">https://github.com/Xe/x/</a></a> , a repo that&#39;s been</p><p>around for over a decade and holds many experiments.</p></li><li><p>The Tigris blog repo: <a href=\"https://github.com/tigrisdata/tigris-blog\" rel=\"nofollow ugc noopener\"><a href=\"https://github.com/tigrisdata/tigris-blog\" rel=\"nofollow ugc noopener\">https://github.com/tigrisdata/tigris-blog</a></a> , a repo that</p><p>has a lot of images and other incompressible objects in it.</p></li></ul>\n<p>Here are the conditions I ran the tests in:</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Item</th><th>Value</th></tr></thead><tbody><tr><td>Started</td><td>2026-09-11T12:28:59-04:00</td></tr><tr><td>Finished</td><td>2026-09-11T13:06:52-04:00</td></tr><tr><td>Host</td><td>Mac, 16 CPUs</td></tr><tr><td>Go</td><td>go1.26.5</td></tr><tr><td>Git</td><td>git version 2.55.0</td></tr><tr><td>Bucket</td><td>xe-objgit-devel</td></tr><tr><td>Build &quot;old&quot;</td><td>v1.0.2</td></tr><tr><td>Build &quot;fixed&quot;</td><td>origin/main</td></tr></tbody></table></div>\n<h3 id=\"push-tests\">Push tests<a href=\"#push-tests\"></a></h3>\n<p>The big thing I wanted to fix was reducing the number of object storage calls. Less object storage calls, less latency waiting for them to resolve. It ended up being ridiculously effective. Object storage call counts sank like a stone.</p>\n<p>All charts in this post are at logarithmic scales so they render more cleanly and the green bars are visible.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Repo</th><th>Build</th><th>Wall</th><th>S3 requests</th><th>PUT</th><th>GET</th><th>HEAD</th><th>LIST</th><th>Keys</th><th>Bucket bytes</th></tr></thead><tbody><tr><td>objgit</td><td>old</td><td>8.7s (8.7s-15.9s)</td><td>231</td><td>46</td><td>47</td><td>0</td><td>138</td><td>42</td><td>829.27 KiB</td></tr><tr><td>objgit</td><td>new</td><td>2.2s (1.9s-2.5s)</td><td>18 (17-20)</td><td>5</td><td>10 (9-12)</td><td>0</td><td>3</td><td>4</td><td>902.98 KiB</td></tr><tr><td>x</td><td>old</td><td>3m29.4s (3m29.4s-3m40.6s)</td><td>9,236 (9,236-9,304)</td><td>1,087</td><td>2,170</td><td>0</td><td>5,979 (5,979-6,047)</td><td>1,082</td><td>54.96 MiB</td></tr><tr><td>x</td><td>new</td><td>14.3s (14.3s-14.7s)</td><td>30</td><td>6</td><td>20</td><td>0</td><td>4</td><td>4</td><td>46.38 MiB</td></tr><tr><td>tigris-blog</td><td>old</td><td>2m13.4s (1m29.1s-2m30.5s)</td><td>3,324 (2,687-3,417)</td><td>515</td><td>522</td><td>0</td><td>2,287 (1,650-2,380)</td><td>511</td><td>354.67 MiB</td></tr><tr><td>tigris-blog</td><td>new</td><td>26.5s (19.2s-27.4s)</td><td>136 (113-146)</td><td>9</td><td>123 (100-133)</td><td>0</td><td>4</td><td>8</td><td>360.75 MiB</td></tr></tbody></table></div>\n<p>The biggest gain was wall clock time for pushing though:</p>\n<p>One of the biggest places that objgit used to lag was pushing taking way longer than it felt like it should. Eliminating the Tigris round trips made pushing way more responsive.</p>\n<h3 id=\"clone-tests\">Clone tests<a href=\"#clone-tests\"></a></h3>\n<p>I also wanted to see how the difference affected clone times. There were the same benefits as with pushing:</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Repo</th><th>Build</th><th>Wall</th><th>S3 requests</th><th>GET</th><th>HEAD</th><th>LIST</th><th>Wire bytes</th></tr></thead><tbody><tr><td>objgit</td><td>old</td><td>11.8s (11.5s-19.5s)</td><td>323</td><td>51</td><td>0</td><td>272</td><td>763.63 KiB</td></tr><tr><td>objgit</td><td>new</td><td>2.6s (1.5s-5.7s)</td><td>17 (16-17)</td><td>13 (12-13)</td><td>0</td><td>4</td><td>767.96 KiB</td></tr><tr><td>x</td><td>old</td><td>3m23.5s (3m23.5s-3m33.6s)</td><td>6,428 (6,428-6,780)</td><td>1,091</td><td>0</td><td>5,337 (5,337-5,689)</td><td>40.68 MiB (40.68 MiB-40.70 MiB)</td></tr><tr><td>x</td><td>new</td><td>54.4s (54.4s-57.8s)</td><td>17 (17-71)</td><td>13 (13-67)</td><td>0</td><td>4</td><td>42.22 MiB (42.22 MiB-42.36 MiB)</td></tr><tr><td>tigris-blog</td><td>old</td><td>2m23.6s (2m19.4s-2m50.2s)</td><td>3,675 (3,601-4,123)</td><td>520</td><td>0</td><td>3,155 (3,081-3,603)</td><td>350.94 MiB (350.93 MiB-351.07 MiB)</td></tr><tr><td>tigris-blog</td><td>new</td><td>1m22s (1m20.8s-1m38.7s)</td><td>158 (137-317)</td><td>155 (134-314)</td><td>0</td><td>3</td><td>349.04 MiB (348.00 MiB-352.67 MiB)</td></tr></tbody></table></div>\n<p>Oh yeah, the time numbers would probably be better if I tested this on a machine with ethernet. I did all my testing with my corp laptop on Wi-Fi to specifically put this in one of the worst conditions it could possibly be in.</p>\n<h2 id=\"conclusion-section\">Conclusion section<a href=\"#conclusion-section\"></a></h2>\n<p>I’m still actively working on this. I’m not confident enough to use this for my\nown projects yet and I wouldn’t blame you for not wanting to use this yet\neither. I still haven’t implemented authentication, authorization, any kind of\nAPI (my long-form <a href=\"https://www.tigrisdata.com/blog/sigv4/\" rel=\"nofollow ugc noopener\">SigV4 auth post</a> was\nactually going to be an objgit post!), or any rate limit beyond what your\nmachine can physically process. If you were to take this, run it, and then\nexpose it to the Internet, then anyone that can connect to that server can pull\nor push whatever they want. Consider not doing that.</p>\n<p>Objgit packfiles also currently accumulate forever, so if you have a bunch of small pushes then there will be a bunch of small packfiles in the bucket. I’m toying with designs that would occasionally compact them into bigger packfiles, but that’s something that can be done later.</p>\n<p>At the least though: Git’s packfile format is a great format for the constraints of storing git repositories in actual filesystems. The moment you put network roundtrips into the mix it all goes south.</p>\n<p>I&#39;m gonna keep working on this and publish reports like this as I learn more. I hope this was interesting! Stay safe out there.</p>\n<p>objgit stores git repositories directly in Tigris. The .bin/.cue container, the columnar index, and the four-tier read ladder in this post are all in the repo.</p>","headings":[{"level":1,"text":"You can run git on object storage if you re-make packfiles","id":"you-can-run-git-on-object-storage-if-you-re-make-packfiles"},{"level":1,"text":"You can run git on object storage if you re-make packfiles","id":"you-can-run-git-on-object-storage-if-you-re-make-packfiles-2"},{"level":2,"text":"Contents","id":"contents"},{"level":2,"text":"What is a Git? A miserable little pile of objects!","id":"what-is-a-git-a-miserable-little-pile-of-objects"},{"level":2,"text":"Yo dawg, herd you like objects","id":"yo-dawg-herd-you-like-objects"},{"level":3,"text":"Messin' with Packfiles","id":"messin-with-packfiles"},{"level":3,"text":"If only you could construct Range requests from packfiles","id":"if-only-you-could-construct-range-requests-from-packfiles"},{"level":2,"text":"SEND CUE SHEET","id":"send-cue-sheet"},{"level":2,"text":"Packfiles v2: object storage boogaloo","id":"packfiles-v2-object-storage-boogaloo"},{"level":3,"text":"Objgit's packfile format that probably needs a name","id":"objgit-s-packfile-format-that-probably-needs-a-name"},{"level":2,"text":"The funny numbers","id":"the-funny-numbers"},{"level":3,"text":"Push tests","id":"push-tests"},{"level":3,"text":"Clone tests","id":"clone-tests"},{"level":2,"text":"Conclusion section","id":"conclusion-section"}]}}