{"article":{"slug":"packing-binary-is-fun-actually","title":"Packing Binary Is Fun, Actually","subtitle":null,"summary":"Pranav Desai’s hands-on tour of binary packing: why packing bits can be fun, the techniques that matter, and practical patterns for packing denser structures without losing your mind.","content_type":"blog_post","language":"en","canonical_url":"https://hereticpleb.vercel.app/blog/packing-binary-is-fun-actually","author":{"name":"Pranav Desai","url":"https://hereticpleb.vercel.app/","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"hereticpleb","url":"https://hereticpleb.vercel.app/","listing_slug":null,"listing":null},"topics":[{"name":"Programming","slug":"programming","url":"https://listedarticles.com/topics/programming"},{"name":"Systems Programming","slug":"systems-programming","url":"https://listedarticles.com/topics/systems-programming"},{"name":"Performance","slug":"performance","url":"https://listedarticles.com/topics/performance"},{"name":"Engineering","slug":"engineering","url":"https://listedarticles.com/topics/engineering"},{"name":"Tutorials","slug":"tutorials","url":"https://listedarticles.com/topics/tutorials"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":5056,"reading_minutes":22,"published_at":"2026-09-24T00:00:00.000Z","added_at":"2026-09-28T09:14:29.303Z","updated_at":"2026-09-28T09:14:29.303Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":false},"profile_url":"https://listedarticles.com/articles/packing-binary-is-fun-actually","markdown_url":"https://listedarticles.com/articles/packing-binary-is-fun-actually.md","example":false,"citation":"Pranav Desai, hereticpleb. \"Packing Binary Is Fun, Actually.\" 24 Sept 2026. https://hereticpleb.vercel.app/blog/packing-binary-is-fun-actually (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://hereticpleb.vercel.app/blog/packing-binary-is-fun-actually"},"body_markdown":"## Why would anyone do this?\n\nI saw someone on twitter arguing that saving data in JSON was apparently not what Real™ developers do.\n\nObviously, I had to become a Real™ Developer too.\n\nTurns out, the answer was binary.\n\nNaturally, I had a brilliant idea:\n\n“How hard could it be to make my own binary format?”\n\nSurely it’s just a little `wb` . ( ˶ˆᗜˆ˵ )\n\nIt was, unfortunately, not just a `wb`.\n\nI ended up building an entire binary schema language that can shrink JSON payloads by 80%.\n\n## What is binary packing?\n\nSay we got some data\n\n`\"hello world\"`\nthen it would be translated in ascii to\n\n```\n104 101 108 108 111    32     119 111 114 108 100\n h   e   l   l   o   [SPACE]   w   o   r   l   d\n```\nSo h becomes 104 in ASCII.\n\nSince these ASCII values fit within 8 bits, each character takes up 1 byte.\n\n```\nh        e        l\n01101000 01100101 01101100 \nl        o        [SPACE]\n01101100 01101111 00100000 \nw        o        r\n01110111 01101111 01110010 \nl        d\n01101100 01100100\n```\nso we can just write it using a lil bitta c:\n\n```\nFILE *f = fopen(\"file.bin\", \"wb\");\nunsigned char data[] = \"hello world\";\nfwrite(data, 1, sizeof(data) - 1, f);\nfclose(f);\n```\nSimple enough. Now let’s try writing 104.\n\nObviously, we could just write 104 as ASCII characters:\n\n`'1' '0' '4' → 49 48 52`\nBut that’s 3 bytes for a number that only needs 1 byte.\n\nSo if we want to save those 2 bytes, we need some way of telling the decoder, “hey, this is an integer, not a string.”\n\nYou could add a header, you would only be adding an additional byte (well depends on how many types you got.. hopefully you don’t have more than 128 types…if you do you got bigger issues.\n\ngreat! lets just use the first byte to represent our type and second to represent our data.\n\n`[TYPE][DATA]`\nsay 0 is int and 1 is string so 104 would be\n\n`00000000 01101000`\nand string would be..\n\n`00000001 01101000`\noh wait…..\n\nthat would only give is `h`\n\nwe need a way to represent different lengths of data. welp lets just get another byte. that should represent the length of our string. so now our binary becomes\n\n`[TYPE][LENGTH][DATA]`\ngreat! now we can represent our string like this:\n\n```\n [TYPE]  [LENGTH] [DATA]\n(STRING)  (11)\n00000001 00001011 00...\nh        e        l\n01101000 01100101 01101100 \nl        o        [SPACE]\n01101100 01101111 00100000 \nw        o        r\n01110111 01101111 01110010 \nl        d\n01101100 01100100\n```\nGREAT! now we could pack both strings and ints together!\nsay we wanted to represent `\"userid\": 123`\n\nnow you could just package it all together\n\n`[TYPE:STRING][LENGTH:6][WORD:userid][TYPE:INT][LENGTH:-][DATA:123]`\nGreat! we can represent 123 as a 1 byte number with 2 bytes of header. but notice, we are not really using LENGTH field for ints? why need it then? waste of bytes eh?\n\nWELL… if we get rid of it, how does our binary reader know where the header ends?\n\nIt needs some way to say “okay, the header is done, start reading the actual data now.”\n\nhuh. what can we use to represent that a byte is ending.\n\nA length byte for the header, perhaps?\n\nEhh. That’s redundant. We’d be removing the length field just to add another length field.\n\nBut hey, we could use a bit in the header itself.\n\nWe could have one bit say:\n\nI am not the last byte in the header. There’s more.\n\nYou might think: why not use the LSB?\n\nWell, then we’d only be able to represent even numbers. Which is… not ideal.\n\nSo we’ll use the MSB instead.\n\nso now our tag looks something like this:\n\n`[CONTINUATION BIT][7 BITS OF DATA]`\nIf the continuation bit is 1, there’s another header byte. If it’s 0, the header is done.\n\nusing this, we can just have our 123 be\n\n`[TYPE=0][DATA=123]`\nand if its a string.\n\n`[TYPE=1][LENGTH=11][DATA=104]...`\nso its of type 1, length 11\n\nbut wait…what if the length is greater than 127? with 7 bits you can only represent up to 127!\n\nWe use the same thing! but for ints!\n\nif the first bit is 1 then the int continues. 128 can be written as:\n\n```\n10000001 00000000\n^\nMSB / continuation bit\n```\n(in big endian)\n\nwhat the binary reader will do:\n\n- Reads the first byte.\n- The MSB is 1, so there’s another byte.\n- The remaining 7 bits are 1.\n- Reads the second byte.\n- Its MSB is 0, so this is the last byte.\n- Its remaining 7 bits are 0.\n- Combines the two 7-bit values to get 128.\n\nThis is a kind of varint (variable-length integer).\n\nThe encoding we’re using here is little-endian: the least-significant 7 bits come first.\n\n`10000000 00000001`\nHey this is great, innit? You can represent different types in the same binary and your binary parser will read them all correctly\n\nBut notice, We are storing this data per field.\n\n```\n[TYPE][DATA]\n[TYPE][DATA]\n[TYPE][DATA]\n```\nAnd most data isn’t just a bunch of random values floating around. It’s usually structured.\n\nTake a C struct:\n\n```\nstruct {\n    int i;\n    char *s;\n    int a[10];\n}\n```\nthis would be say on a 32 bit system.\n\n`[32-bit int] [32-bit pointer] [10 × 32-bit ints]`\nand we didn’t have to add headers everytime. because we know the type of the data from the struct itself.\n\nHmm. I wonder if we can do this for our binary data…\n\nAnd yes, we can.\n\nThat’s what a schema is!\n\nso for our struct our schema can just be:\n\n```\ni: int\ns: char *\na: list(int)\n```\nThe schema lets us know the type without storing the type alongside every value.\n\nnow our binary format doesn’t need to worry about the type! it only need to worry the size of the data!\n\nThat’s what protobuf does\n\nSo lets think about all the different sizes of data we can have.\n\nwe got ints, we got floats, bools, strings.\n\nWe can treat ints and bools as varints, while floats are fixed-width: f32 or f64.\n\nStrings are different. We can’t just encode their bytes as a varint, because the bytes themselves are the actual data we need to preserve.\n\nSo instead, we need to know how many bytes belong to the string before we start reading it.\n\nso now our encoding types are:\n\n- varint\n- f32\n- f64\n- delimited\n\nbut wait! how do we access our fields? like we cant just go “gimme string” we will need to spacify which one. and no problem lets just represent each one with a number. we can index them or have the user assign their own numbers to address them. this is what we need field number for.\n\nAnd notice something else: we only have four possible types.\n\nFour values fit perfectly into 2 bits.\n\n```\n00 - varint \n01 - f32\n10 - f64\n11 - delimited\n```\nhey isn’t that neat. now what we could do is just encode it…*inside our field number!*\n\nwait how?\n\nsome binary trickery..not really.\n\nYou just shove them together.\n\nMove the field number left by 2 bits to make room for our 2-bit type, then OR the type into those empty bits.\n\nsay your field number is 10 and it is a delimited type.\n\n```\n00001010 (10)\nleft shift that by 2\n00101000\nOR it with our type!\n00101011\n```\nLOOK AT THAT! our 1 byte number tells us both what type it is and what its field number is!\n\nBut what if we go FURTHER.\n\nWe’ve already packed the type into the field number.\n\nWhy stop there?\n\nWhat if we could pack the length in too?\n\nWe can.\n\nnow our header can hold\n\n`[FIELD NUMBER][LENGTH][TYPE]`\nALL in a singular varint!!!! ◝(ᵔᗜᵔ)◜\n\nCool. Except for one thing\n\nits just so much tedious work. To pack a string, I have to tell it ‘this is delimited,’ give it the length, and then give it the bytes. Every. Single. Time.\n\nPEASANTS DO THAT. Plebeians. we don’t do that.\n\nSo obviously, the solution is to write an entire schema language. We declare our data in a schema file, and let the program deal with all that tedious formatting and encoding nonsense.\n\n## Building the schema syntax\n\nalright soo… we need to decide on a syntax that doesn’t suck your soul (looking at you protobuf)\n\nfield numbers.. what are they? index right. how do you index stuff in your grocery list? you write number. item why not use the same!\n\nso something like this:\n\n`1. name`\nwe need the schema to represent the types. lets just steal how other languages do it and do it like this:\n\n`1. name: type`\nneat huh. but wait. how do we indicate the end of a message(struct)\n\nwell we could do {} but its not very nice is it. why over complicate stuff its a list. lists have an end. lets have an end.\n\n```\nmessage name:\n    1. name:type\n    2. name:type\n    3. name:type\nend\n```\nand no, indentations shouldn’t matter. its so annoying to work with languages where indentation matters. its just painful. lets just not do that.\n\nso..how will you identify the end of the line without a semicolon? new line char?\n\nI mean we could do that but what if they wanted to type it in a single line, its ugly but say they want to for whatever reason. lets not restrict that. but semicolons are ugly.\n\nwe could use the number! if we see a number and a dot we just consider it a new entry!\n\nneat huh.𐔌*ˊᵕˋ*𐦯\n\nproblem…what if the field is just not found when the compiler is reading it? we could crash…and we would for all the fields but we do want optional fields don’t we. lets just go with the obvious route and do something like this:\n\n    `number. name: type = defaultValue`\nand that’s optional. pretty intuitive.\n\nnow that we are here anyway lets think about all the different kinds of data people could represent…\n\nwell we obviously got our entries of types.\n\noh we would need lists. we need some sort of way to pick between a few things so we would need an enum. great lets think about those..\n\noh we could just have them be function like, that’s pretty intuitive.\n\n    `number. name: structure(type)`\nlike:\n\n    `1. friends: list(People)`\noh wait but what is people? oh it should be a message too. something like this:\n\n```\nmessage People:\n    1. name: string\n    2. location: string\nend\n```\nhuh.. we need custom types as well… so in the language do we want to have everything be in order? ehh that’s pretty cringe later we could just do a pass and put together all of our table and just have it refer to that struct.\n\nhey thanks to this we can have self reference as well. cuz when we pack it it will just be a pointer to the struct! so we can do something like:\n\n```\nmessage People:\n    1. name: string\n    2. location: string\n    3. friends: list(People)\n```\nnow for enums. cuz we have two passes we can just put enums outside of message! it doesn’t matter where it is declared either! we can declare our enum something like this:\n\n```\nenum Name:\n    number. name \n    number. name\n    number. name\nend\n```\nalso protobuf does this thing where it forces you to have the enum start from 0 like do we REALLY need that? is the compiler so dumb it cant tell? lets just have the compiler handle that by default\n\noh wait…. people might need some way to represent an enum but all the types are not the same. yup. that’s a union. lets just add that no problemo\n\n    `number. name: union(type, type, type...)`\noh wait.. our default value. how would that work with unions. if we have a union like this:\n\n    `number. name: union(i32, i64) = 10`\nit is ambiguous weather 10 is an i32 or an i64 so when field is not found what should we do?\n\nwe could fix this by having the default be next to the type.\n\n    `number. name: union(i32 = 10, i64)`\nnow if a the field is not found it will be an i32 with value 10\n\nokay well people would wanna map stuff as well…so we add\n\n    `number. name: map(key_type, value_type)`\nlets just add syntax for declaring packages and importing stuff:\n\n`package \"package name\"``import \"package\"`\nhey would you look at that we have some very neat syntax. lets just put it all together:\n\n```\npackage \"com.game.core\"\nimport \"math.jbin\"\nenum Activity:\n    1. active\n    2. inactive\nend\nmessage Player:\n    1. name: string\n    2. health: i32 = 100\n    3. weapons: list(string)\n    4. connections: list(Player)\n    5. activeStatus: Activity\n    6. inventory: map(string, i32)\n    7. balance: union(string=\"empty\", i32)\nend\n```\noh wait…we have a whole language now….\n\nanyway. this is the equivalent of it in .proto\n\n```\nsyntax = \"proto2\";\npackage com.game.core;\nimport \"math.proto\";\nenum Activity {\n  ACTIVE = 1;\n  INACTIVE = 2;\n}\nmessage Player {\n  optional string name = 1;\n  optional int32 health = 2 [default = 100];\n  \n  repeated string weapons = 3;\n  repeated Player connections = 4;\n  \n  optional Activity active_status = 5;\n  map<string, int32> inventory = 6;\n  \n  oneof balance {\n    string empty = 7;\n    int32 amount = 8;\n  }\n}\n```\nlook at that. ew.\n\n## Time to actually implement this\n\nTo implement this I naturally chose C++.\n\nBy “naturally,” I mean I wanted something I could put on a resume and wasn’t in the mood to fight the Rust borrow checker. I have aged enough.\n\nOkay. Enough designing. Time to actually make the thing work.\n\nFirst, we need to encode our binary.\n\n### Encoding and Decoding our Data\n\nTo encode a varint, it’s basically just this:\n\n```\nvoid encodeVariant(std::vector<uint8_t> &buffer, uint64_t value) {\n    while (value >= 128) {\n        uint8_t lower = value & 127; // 127 = 01111111\n        lower |= 128;\n        buffer.push_back(lower);\n        value >>= 7;\n    }\n    uint8_t lower = value & 127;\n    buffer.push_back(lower);\n}\n```\nThis builds our varint. If the value is >= 128, we take the lowest 7 bits, set the MSB to 1 to say “there’s more,” and append it to the buffer.\n\nThen we shift the value by 7 bits and repeat.\n\nFor the final byte, we leave the MSB at 0.\n\npretty simple. and we can decode it using this:\n\n```\nuint64_t decodeVariant(const std::vector<uint8_t> &buffer, size_t &offset) {\n    uint64_t result = 0;\n    int shift = 0;\n    while ((buffer[offset] & 128) != 0 && offset < buffer.size()) {\n        uint8_t byte = buffer[offset++];\n        byte &= 127;\n        const uint64_t cast = static_cast<uint64_t>(byte) << shift;\n        result |= cast;\n        shift += 7;\n    }\n    uint8_t byte = buffer[offset++] & 127;\n    const uint64_t cast = static_cast<uint64_t>(byte) << shift;\n    result |= cast;\n    return result;\n}\n```\nNow that we can encode and decode our varints using LEB128 lets encode and decode some tags(our header metadata we discussed about)\n\n```\nvoid encodeTag(std::vector<uint8_t> &buffer, uint32_t fieldNumber,\n               wiretype wiretype) {\n    uint64_t val = (static_cast<uint64_t>(fieldNumber) << 2);\n    val |= static_cast<uint64_t>(wiretype);\n    encodeVariant(buffer, val);\n}\nvoid decodeTag(const std::vector<uint8_t> &buffer, size_t &offset,\n               uint32_t &outFieldNumber, wiretype &outType) {\n    uint64_t val = decodeVariant(buffer, offset);\n    outType = static_cast<wiretype>(val & 3);\n    outFieldNumber = val >> 2;\n}\n```\nAnd we can add a couple of helpers for strings:\n\n```\nvoid encodeString(std::vector<uint8_t> &buffer, uint32_t fieldNumber,\n                  const std::string &text) {\n    encodeTag(buffer, fieldNumber, wiretype::Delimited);\n    encodeVariant(buffer, text.size());\n    buffer.insert(buffer.end(), text.begin(), text.end());\n}\nstd::string decodeString(const std::vector<uint8_t> &buffer, size_t &offset) {\n    uint64_t size = decodeVariant(buffer, offset);\n    std::string result =\n        std::string(buffer.begin() + offset, buffer.begin() + offset + size);\n    offset += size;\n    return result;\n}\n```\nEncoding a string is now just three things: write its tag, write its length, then write its bytes.\n\nbelieve it or not, that’s basically all our core engine done! the rest is just the compiler and the json conversion stuff! ٩(^ᗜ^ )و ´-\n\nGreat! now that we have written our binary encoding… lets make a compiler should be simple…right?\n\n### Building Compiler\n\nYes it is a compiler. stop it I don’t wanna call it a transpiler it is a compiler. the definition is:\n\ncompiler is a computer program that translates source code written in one programming language (the source language) into another programming language (the target language), while preserving the exact meaning and behavior of the original code.\n\nours does that. gtfo compiler people. my program takes my schema and translates it into either a binary or code.\n\n#### The problem\n\nOur language is pretty cute. Pretty slick.\n\nUnfortunately, the computer has no idea what any of it means.\n\nIt’s just bytes.\n\nso lets assign meaning to these symbols!\n\nlets take our syntax:\n\n    `1. name:string`\nso its in the structure of:\n\n`number → dot → identifier → colon → type`\nand then we can use this to build our representation of these entries. that’s what a Lexer does\n\n#### Building the Lexer\n\nYes lexer is like what you think it is. its just a big loop with a bunch of if statements. All it does is read our file and spit out these tokens.(not the ai kind)\n\n```\nenum class TokenType {\n    Keyword_Message,\n    Keyword_Enum,\n    Keyword_Optional,\n    Keyword_Map,\n    Keyword_Union,\n    Keyword_End,\n    Keyword_Package,\n    Keyword_Import,\n    Identifier,\n    Number,\n    StringLiteral,\n    Comment,\n    Equals,\n    Colon,\n    Comma,\n    Dot,\n    LParen,\n    RParen,\n    EndOfFile\n};\n```\nso our file\n\n```\nmessage User:\n    1. name: string\n    2. id: i32\n```\nthe tokenized output would be:\n\n```\n    [Keyword_Message]\n    [Identifier: \"User\"]\n    [Colon]\n    [Number: 1]\n    [Dot]\n    [Identifier: \"name\"]\n    [Colon]\n    [Identifier: \"string\"]\n    [Number: 2]\n    [Dot]\n    [Identifier: \"id\"]\n    [Colon]\n    [Identifier: \"i32\"]\n    [EndOfFile]\n```\nPretty neat. Now the parser doesn’t have to decipher a bunch of raw characters. It can just work with these tokens. parser looks at these and then actually emits the AST (abstract syntax tree).\n\n#### What is AST?\n\nIts just how we represent our program data in graphLang it was a literal tree node that held all the data and all data was just a single shape. For this one we can define our schema as the root. we just have two kinds of messages rn messages and enums so our head can be defined as:\n\n```\nstruct Schema {\n    std::vector<EnumDef> enums;\n    std::vector<MessageDef> messages;\n};\n```\nso our Schema will be the root node and the tree will look like this:\n\n```\n    (Schema)\n    /      \\\n(enums) (messages)\n```\n#### Building the AST\n\nsoo…. what is enums and messages? from the definition above you can see its a vector of structs. our messages can be defined as:\n\n```\nstruct MessageDef {\n    std::string name;\n    std::vector<Field> fields;\n    int line = 0;\n    std::string comment = \"\";\n};\n```\nand then our Field as:\n\n```\nstruct Field {\n    uint32_t number;\n    std::string name;\n    DataType type;\n    bool isOptional = false;\n    std::string defaultValue = \"\";\n    int line = 0;\n    std::string comment = \"\";\n};\n```\nnow our AST looks like this:\n\n```\nSchema\n└── messages\n    └── MessageDef\n        ├── name\n        ├── line\n        ├── comment\n        └── fields\n            └── Field\n                ├── number\n                ├── name\n                ├── type\n                ├── isOptional\n                ├── defaultValue\n                ├── line\n                └── comment\n```\nand all that’s left is our enum definition:\n\n```\nstruct EnumDef {\n    std::string name;\n    std::vector<EnumEntry> entries;\n    int line = 0;\n    std::string comment = \"\";\n};\n```\nand our enum entry:\n\n```\nstruct EnumEntry {\n    uint32_t number;\n    std::string name;\n    int line = 0;\n    std::string comment = \"\";\n};\n```\nGreat! we got all our pieces. our AST looks like this now:\n\n```\nSchema\n├── Enums\n│   └── EnumDef\n│       ├── name\n│       └── entries\n│           └── EnumEntry\n│               ├── number\n│               └── name\n│\n└── Messages\n    └── MessageDef\n        ├── name\n        └── fields\n            └── Field\n                ├── number\n                ├── name\n                ├── type\n                ├── isOptional\n                └── defaultValue\n```\n#### Building the parser\n\nwe use this and build our parser. and yes. our parser is just a loop with if statements (well technically recursive decent or whatever but recursion is just a form of iteration).\n\nit is actually pretty simple the whole loop:\n\n```\nSchema Parser::parse() {\n    Schema s;\n    while (!isAtEnd()) {\n        if (peek().type == TokenType::Comment) {\n            consume();\n            continue;\n        }\n        if (peek().type == TokenType::Keyword_Package) {\n            consume();\n            if (peek().type == TokenType::StringLiteral) {\n                s.packageName = consume().value;\n            } else {\n                error(\"Expected string literal after package\");\n            }\n        } else if (peek().type == TokenType::Keyword_Import) {\n            consume();\n            if (peek().type == TokenType::StringLiteral) {\n                s.imports.push_back(consume().value);\n            } else {\n                error(\"Expected string literal after import\");\n            }\n        } else if (peek().type == TokenType::Keyword_Message)\n            s.messages.push_back(parseMessage());\n        else if (peek().type == TokenType::Keyword_Enum)\n            s.enums.push_back(parseEnum());\n        else\n            error(\"Unexpected token in global scope\");\n    }\n    return s;\n}\n```\nand for each of the types its just a bunch if conditions checking and then returning the AST node.\n`consume()` returns the current token and moves the pointer over by one and `peek()` gives you the next token data without moving the pointer\n\nGREAT! lets test it out.\n\n```\n    message User:\n        1. name: adfasdfasdfasdfasd\n        2. id: 012349\n```\nlets see what our parser says:\n\n“Mighty good mate! seems excellent innit? want a cuppa?”\n\nyeah..so we need a thing that checks the user isn’t just syntactically correct but also the shit makes sense.\n\nso we need a fact checker, a twitter community note if you will.\n\nAnd that’s what our semantic analyzer does!\n\n#### Building the Semantic Analyzer\n\nSo to validate stuff we will be doing it in two passes.\n\n- First pass we build all of our symbols and put them into a table. (Basically just all the valid stuff that are identifiable)\n- Second pass we use that the table we built validate the AST.\n\n```\nvoid SemanticAnalyzer::analyze(const Schema &schema) {\n    buildSymbolTable(schema);\n    validateEnums(schema);\n    validateMessages(schema);\n    validateCyclicDependencies(schema);\n}\n```\nBecause we do this in two passes, declaration order doesn’t matter. Which is pretty neat.\n\nSo during `validateMessages` , when it looks at\n`adfasdfasdfasdfasd` , it checks our symbol table,\nrealizes that type doesn’t exist, and throws an error . It also checks\nthat you didn’t do something stupid like use field\nnumber  1  twice, or use a list as a map key.\n\nGreat! Look at what we got so far!\n\n- A Lexer that chops everything up into tokens\n- A Parser that takes the tokens and builds the AST\n- A Semantic Analyzer that validates the AST and throws errors\n- An encoder/decoder using LEB128.\n\nNow all that’s left to do is to do two things.\n\n- Convert JSON files into binary dynamically\n- Codegen from the schema so they can use them in their language natively\n\n### Dynamic Packer\n\nNow for the dynamic packer. It takes a schema, takes a JSON file, and builds our binary.\n\nFor JSON, I just used `nlohmann/json`. Parsing JSON on top of everything else would’ve been a completely unnecessary side quest.\n\nSo we have our JSON loaded in memory. We have our AST loaded in memory. Now we just walk them together.\n\nWhen the JSON parser sees the key “id”: 123 , it doesn’t know what to do. But it asks the AST! The AST says,\n\n“Oh, id ? That’s field number 2, and it’s an i32 .”\n\nSo the Packer just calls our  `encodeTag()`  and  `encodeVariant()`\nfunctions and spits the bytes into our buffer.\n\nNotice what we didn’t do? We didn’t generate any C++ code. We didn’t compile any wrappers. We just read the JSON, read the schema, and built the binary dynamically.\n\ntake that protobuf\n\n**Lets test it out!**\n\n```\n{\n    \"name\": \"Pranav\",\n    \"health\": 100\n}\n```\nIf we save this as JSON, with spaces and quotes, it’s about 35 bytes\n\nlets pack it in binary.\n\nand that is…\n\n**NINE BYTES**\n\n**hell yeahh look at that!!!**\n\n74.29% REDUCTION\n\nit would be even higher if we didn’t have strings and stuff. 6 of the 9 bytes are used for the word “Pranav” but besides the point.\n\nwe have reduced our size of storage by a LOT.\n\nAnd decoding is way simpler too. The binary already tells us what each field is supposed to be, so the decoder doesn’t have to deal with JSON’s syntax and type representation.\n\nBut what if you don’t want to pay the cost of dynamically looking everything up at runtime? What if you just want normal structs in your language?\n\nThat’s where AOT code generation comes in.\n\n### Codegen\n\nHmm… how would one generate code? I mean we have our AST so we know what it looks like and what it is semantically so like its just translating that into our language..\n\nyou could write a `cCodeGen()` func and add it to our class and call.\n\nif we wanted a generator for python? oh that’s easy! i’ll just add a `pythonCodeGen`!\n\noh I need a `jsCodeGen()` okay look. too far. why are you writing javascript. but regardless.\n\nBut apparently the customer is always right in their language preferences or whatever. okay now this class is getting a wee bit too bloated for my liking.\n\nThat’s exactly why we need a visitor pattern!\n\n#### Visitor Pattern\n\nSo what is a visitor? in hindsight, design pattern cope for ones who’s language doesn not have algebraic datatypes. So why did I use it even though cpp has the `std::variant`? Idk I read it in an article once. I wanted ot learn about it.\n\nAnyway. so a visitor pattern is just a lil handshake the caller and callee do. we get type safety from it. in our implementation our caller passes itself in and becomes the visitor. the callee or the acceptor accepts the visitor and does a little func call on the visitor passing itself in triggering a double dispatch. pretty neat..but its just so much mental overhead for a simple problem.\n\nto accomplish this we will need to edit our structs in the ADT to also hold a function.\n\n`void accept(SchemaVisitor &visitor) const;`\nand we create an abstract class with a bunch of virtual methods that our generators implement:\n\n```\n#pragma once\n#include \"Schema.hpp\"\nclass SchemaVisitor {\n  public:\n    virtual ~SchemaVisitor() = default;\n    virtual void visit(const Schema &s) = 0;\n    virtual void visit(const MessageDef &m) = 0;\n    virtual void visit(const EnumDef &e) = 0;\n    virtual void visit(const EnumEntry &ee) = 0;\n    virtual void visit(const Field &f) = 0;\n};\n```\nso our c generator for example looks like this:\n\n```\n#pragma once\n#include \"SchemaVisitor.hpp\"\n#include <ostream>\n#include <string>\nclass CGenerator : public SchemaVisitor {\n  private:\n    std::ostream &out;\n    std::string currentEnumName;\n    const Schema* currentSchema = nullptr;\n  public:\n    CGenerator(std::ostream &outputStream) : out(outputStream) {}\n    void visit(const Schema &schema) override;\n    void visit(const MessageDef &message) override;\n    void visit(const EnumDef &enumDef) override;\n    void visit(const EnumEntry &ee) override;\n    void visit(const Field &field) override;\n};\n```\nso our C generator just overrides these funcs and in these visits it does this:\n\n```\nvoid CGenerator::visit(const EnumDef &enumDef) {\n    currentEnumName = enumDef.name;\n    out << \"typedef enum {\\n\";\n    for (const auto &ee : enumDef.entries) {\n        ee.accept(*this);\n    }\n    out << \"} \" << enumDef.name << \";\\n\\n\";\n}\n```\nwhen we do ee.accept(*this) it does this:\n\n`void EnumEntry::accept(SchemaVisitor &visitor) const { visitor.visit(*this); }`\nso it just calls CGenerator.visit(ee) so the visit for environment entries is called in CGenerator.\n\n```\nvoid CGenerator::visit(const EnumEntry &ee) {\n    out << \"    \" << currentEnumName << \"_\" << ee.name << \" = \" << ee.number << \",\\n\";\n}\n```\nand this is done for each type. and that’s how the code is generated.\n\nGREAT! lets test it out..\n\noh wait. we cant. we don’t have a way to do that. we need a cli.\n\n### Building CLI\n\nFor the CLI I used jarro2783/cxxopts cuz I didn’t want to do manual parsing. it will be a fun project both this and the json I will do them sometime else but for now I used these.\n\nAnd building the cli was pretty easy from this library you get a bunch of stuff that just works.\n\nI just wrote my options.\n\n```\noptions.add_options()(\"command\", \"Command to run (e.g. build, pack)\",\n                              cxxopts::value<std::string>())(\n            \"input\", \"Input schema file\", cxxopts::value<std::string>())(\n            \"o,out\",\n            \"Output (target language for build, or output binary file for \"\n            \"pack)\",\n            cxxopts::value<std::string>())(\"j,json\",\n                                           \"Input JSON file (for pack command)\",\n                                           cxxopts::value<std::string>())(\n            \"m,msg\", \"Root message name to pack (for pack command)\",\n            cxxopts::value<std::string>())(\"h,help\", \"Print usage\");\n        options.parse_positional({\"command\", \"input\"});\n        auto result = options.parse(argc, argv);\n```\nand it works. it was great. and these options were handled in a bunch of if statements(now that I think about it most of this project has just been a loop and a bunch of if statements)\n\n```\nif (command == \"build\") {\n            if (!result.count(\"input\")) {\n                std::cerr << \"Error: No input file specified.\" << std::endl;\n                return 1;\n            }\n            if (!result.count(\"out\")) {\n                std::cerr << \"Error: --out flag is required (e.g., --out c).\"\n                          << std::endl;\n                return 1;\n            }\n            std::string inputFile = result[\"input\"].as<std::string>();\n            std::string targetLang = result[\"out\"].as<std::string>();\n            std::ifstream file(inputFile);\n            if (!file.is_open()) {\n                std::cerr << \"Error: Could not open file \" << inputFile\n                          << std::endl;\n                return 1;\n            }\n            std::stringstream buffer;\n            buffer << file.rdbuf();\n            std::string schemaText = buffer.str();\n            std::vector<Token> tokens = tokenize(schemaText);\n            Parser parser(tokens);\n            Schema schema = parser.parse();\n            SemanticAnalyzer analyzer;\n            analyzer.analyze(schema);\n            if (targetLang == \"c\") {\n                CGenerator cGen(std::cout);\n                schema.accept(cGen);\n            } else if (targetLang == \"py\" || targetLang == \"python\") {\n                PythonGenerator pyGen(std::cout);\n                schema.accept(pyGen);\n            } else {\n                std::cerr << \"Code generation for '\" << targetLang\n                          << \"' is not supported yet!\" << std::endl;\n            }\n```\nanyhow. lets test it out.\n\nwe will write our schema as this:\n\n```\npackage \"com.mmo.game\"\nenum Faction:\n    1. alliance\n    2. horde\n    3. neutral\nend\nmessage Vector3:\n    1. x: i32\n    2. y: i32\n    3. z: i32\nend\nmessage InventoryItem:\n    1. itemId: i32\n    2. quantity: i32\n    3. isSoulbound: bool\nend\nmessage Character:\n    1. id: i64\n    2. name: string\n    3. level: i32\n    4. faction: Faction\n    5. position: Vector3\n    6. inventory: list(InventoryItem)\n    7. attributes: map(string, i32)\nend\n```\nand run build with output as c… and..\n\n```\n#pragma once\n#include <stdint.h>\n#include <stdbool.h>\n#include <string.h>\n#include <stdlib.h>\ntypedef struct Vector3 Vector3;\ntypedef struct InventoryItem InventoryItem;\ntypedef struct Character Character;\ntypedef struct {\n    InventoryItem* data;\n    size_t length;\n    size_t capacity;\n} jbin_list_InventoryItem;\ntypedef struct {\n    char** keys;\n    int32_t* values;\n    size_t length;\n    size_t capacity;\n} jbin_map_char_ptr_int32;\ntypedef enum {\n    Faction_alliance = 1,\n    Faction_horde = 2,\n    Faction_neutral = 3,\n} Faction;\nstruct Vector3 {\n    int32_t x;\n    int32_t y;\n    int32_t z;\n};\nstruct InventoryItem {\n    int32_t itemId;\n    int32_t quantity;\n    bool isSoulbound;\n};\nstruct Character {\n    int64_t id;\n    char* name;\n    int32_t level;\n    Faction faction;\n    Vector3 position;\n    jbin_list_InventoryItem inventory;\n    jbin_map_char_ptr_int32 attributes;\n};\n```\nLOOK AT THAT!! zero dependency left from the program. we provide all the stuff it needs from our schema!!!\n\n(note: that isn’t the full c file that was generated. there were also a lot of pack/unpack functions for individual messages, setters and getters and encode/decode funcs)\n\nwe can also encode to binary by passing in a json file.\n\n```\n{\n    \"id\": 123456789,\n    \"name\": \"LeroyJenkins\",\n    \"level\": 60,\n    \"faction\": \"alliance\",\n    \"position\": {\n        \"x\": 100,\n        \"y\": 200,\n        \"z\": 300\n    },\n    \"inventory\": [\n        { \"itemId\": 999, \"quantity\": 1, \"isSoulbound\": true },\n        { \"itemId\": 45, \"quantity\": 100, \"isSoulbound\": false }\n    ],\n    \"attributes\": {\n        \"strength\": 120,\n        \"agility\": 45\n    }\n}\n```\nand the size of this json is `401 bytes`\n\nlets pack it into binary. and the file size is..\n\n`80 bytes!`\n\nEIGHTY PERCENT REDUCTION\n\n`80.05%`\n\n## Conclusion\n\nSo yeah. I started out just wanting to save some JSON out of spite, and I accidentally built a lexer, a parser, an AST, a semantic analyzer, a dynamic binary packer, and a multi-language code generator. And honestly? Packing binary is fun, actually.\n\nanyhow, checkout the repo.","body_html":"<h2 id=\"why-would-anyone-do-this\">Why would anyone do this?</h2>\n<p>I saw someone on twitter arguing that saving data in JSON was apparently not what Real™ developers do.</p>\n<p>Obviously, I had to become a Real™ Developer too.</p>\n<p>Turns out, the answer was binary.</p>\n<p>Naturally, I had a brilliant idea:</p>\n<p>“How hard could it be to make my own binary format?”</p>\n<p>Surely it’s just a little <code>wb</code> . ( ˶ˆᗜˆ˵ )</p>\n<p>It was, unfortunately, not just a <code>wb</code>.</p>\n<p>I ended up building an entire binary schema language that can shrink JSON payloads by 80%.</p>\n<h2 id=\"what-is-binary-packing\">What is binary packing?</h2>\n<p>Say we got some data</p>\n<p><code>&quot;hello world&quot;</code>\nthen it would be translated in ascii to</p>\n<pre><code>104 101 108 108 111    32     119 111 114 108 100\n h   e   l   l   o   [SPACE]   w   o   r   l   d</code></pre>\n<p>So h becomes 104 in ASCII.</p>\n<p>Since these ASCII values fit within 8 bits, each character takes up 1 byte.</p>\n<pre><code>h        e        l\n01101000 01100101 01101100 \nl        o        [SPACE]\n01101100 01101111 00100000 \nw        o        r\n01110111 01101111 01110010 \nl        d\n01101100 01100100</code></pre>\n<p>so we can just write it using a lil bitta c:</p>\n<pre><code>FILE *f = fopen(&quot;file.bin&quot;, &quot;wb&quot;);\nunsigned char data[] = &quot;hello world&quot;;\nfwrite(data, 1, sizeof(data) - 1, f);\nfclose(f);</code></pre>\n<p>Simple enough. Now let’s try writing 104.</p>\n<p>Obviously, we could just write 104 as ASCII characters:</p>\n<p><code>&#39;1&#39; &#39;0&#39; &#39;4&#39; → 49 48 52</code>\nBut that’s 3 bytes for a number that only needs 1 byte.</p>\n<p>So if we want to save those 2 bytes, we need some way of telling the decoder, “hey, this is an integer, not a string.”</p>\n<p>You could add a header, you would only be adding an additional byte (well depends on how many types you got.. hopefully you don’t have more than 128 types…if you do you got bigger issues.</p>\n<p>great! lets just use the first byte to represent our type and second to represent our data.</p>\n<p><code>[TYPE][DATA]</code>\nsay 0 is int and 1 is string so 104 would be</p>\n<p><code>00000000 01101000</code>\nand string would be..</p>\n<p><code>00000001 01101000</code>\noh wait…..</p>\n<p>that would only give is <code>h</code></p>\n<p>we need a way to represent different lengths of data. welp lets just get another byte. that should represent the length of our string. so now our binary becomes</p>\n<p><code>[TYPE][LENGTH][DATA]</code>\ngreat! now we can represent our string like this:</p>\n<pre><code> [TYPE]  [LENGTH] [DATA]\n(STRING)  (11)\n00000001 00001011 00...\nh        e        l\n01101000 01100101 01101100 \nl        o        [SPACE]\n01101100 01101111 00100000 \nw        o        r\n01110111 01101111 01110010 \nl        d\n01101100 01100100</code></pre>\n<p>GREAT! now we could pack both strings and ints together!\nsay we wanted to represent <code>&quot;userid&quot;: 123</code></p>\n<p>now you could just package it all together</p>\n<p><code>[TYPE:STRING][LENGTH:6][WORD:userid][TYPE:INT][LENGTH:-][DATA:123]</code>\nGreat! we can represent 123 as a 1 byte number with 2 bytes of header. but notice, we are not really using LENGTH field for ints? why need it then? waste of bytes eh?</p>\n<p>WELL… if we get rid of it, how does our binary reader know where the header ends?</p>\n<p>It needs some way to say “okay, the header is done, start reading the actual data now.”</p>\n<p>huh. what can we use to represent that a byte is ending.</p>\n<p>A length byte for the header, perhaps?</p>\n<p>Ehh. That’s redundant. We’d be removing the length field just to add another length field.</p>\n<p>But hey, we could use a bit in the header itself.</p>\n<p>We could have one bit say:</p>\n<p>I am not the last byte in the header. There’s more.</p>\n<p>You might think: why not use the LSB?</p>\n<p>Well, then we’d only be able to represent even numbers. Which is… not ideal.</p>\n<p>So we’ll use the MSB instead.</p>\n<p>so now our tag looks something like this:</p>\n<p><code>[CONTINUATION BIT][7 BITS OF DATA]</code>\nIf the continuation bit is 1, there’s another header byte. If it’s 0, the header is done.</p>\n<p>using this, we can just have our 123 be</p>\n<p><code>[TYPE=0][DATA=123]</code>\nand if its a string.</p>\n<p><code>[TYPE=1][LENGTH=11][DATA=104]...</code>\nso its of type 1, length 11</p>\n<p>but wait…what if the length is greater than 127? with 7 bits you can only represent up to 127!</p>\n<p>We use the same thing! but for ints!</p>\n<p>if the first bit is 1 then the int continues. 128 can be written as:</p>\n<pre><code>10000001 00000000\n^\nMSB / continuation bit</code></pre>\n<p>(in big endian)</p>\n<p>what the binary reader will do:</p>\n<ul><li>Reads the first byte.</li><li>The MSB is 1, so there’s another byte.</li><li>The remaining 7 bits are 1.</li><li>Reads the second byte.</li><li>Its MSB is 0, so this is the last byte.</li><li>Its remaining 7 bits are 0.</li><li>Combines the two 7-bit values to get 128.</li></ul>\n<p>This is a kind of varint (variable-length integer).</p>\n<p>The encoding we’re using here is little-endian: the least-significant 7 bits come first.</p>\n<p><code>10000000 00000001</code>\nHey this is great, innit? You can represent different types in the same binary and your binary parser will read them all correctly</p>\n<p>But notice, We are storing this data per field.</p>\n<pre><code>[TYPE][DATA]\n[TYPE][DATA]\n[TYPE][DATA]</code></pre>\n<p>And most data isn’t just a bunch of random values floating around. It’s usually structured.</p>\n<p>Take a C struct:</p>\n<pre><code>struct {\n    int i;\n    char *s;\n    int a[10];\n}</code></pre>\n<p>this would be say on a 32 bit system.</p>\n<p><code>[32-bit int] [32-bit pointer] [10 × 32-bit ints]</code>\nand we didn’t have to add headers everytime. because we know the type of the data from the struct itself.</p>\n<p>Hmm. I wonder if we can do this for our binary data…</p>\n<p>And yes, we can.</p>\n<p>That’s what a schema is!</p>\n<p>so for our struct our schema can just be:</p>\n<pre><code>i: int\ns: char *\na: list(int)</code></pre>\n<p>The schema lets us know the type without storing the type alongside every value.</p>\n<p>now our binary format doesn’t need to worry about the type! it only need to worry the size of the data!</p>\n<p>That’s what protobuf does</p>\n<p>So lets think about all the different sizes of data we can have.</p>\n<p>we got ints, we got floats, bools, strings.</p>\n<p>We can treat ints and bools as varints, while floats are fixed-width: f32 or f64.</p>\n<p>Strings are different. We can’t just encode their bytes as a varint, because the bytes themselves are the actual data we need to preserve.</p>\n<p>So instead, we need to know how many bytes belong to the string before we start reading it.</p>\n<p>so now our encoding types are:</p>\n<ul><li>varint</li><li>f32</li><li>f64</li><li>delimited</li></ul>\n<p>but wait! how do we access our fields? like we cant just go “gimme string” we will need to spacify which one. and no problem lets just represent each one with a number. we can index them or have the user assign their own numbers to address them. this is what we need field number for.</p>\n<p>And notice something else: we only have four possible types.</p>\n<p>Four values fit perfectly into 2 bits.</p>\n<pre><code>00 - varint \n01 - f32\n10 - f64\n11 - delimited</code></pre>\n<p>hey isn’t that neat. now what we could do is just encode it…<em>inside our field number!</em></p>\n<p>wait how?</p>\n<p>some binary trickery..not really.</p>\n<p>You just shove them together.</p>\n<p>Move the field number left by 2 bits to make room for our 2-bit type, then OR the type into those empty bits.</p>\n<p>say your field number is 10 and it is a delimited type.</p>\n<pre><code>00001010 (10)\nleft shift that by 2\n00101000\nOR it with our type!\n00101011</code></pre>\n<p>LOOK AT THAT! our 1 byte number tells us both what type it is and what its field number is!</p>\n<p>But what if we go FURTHER.</p>\n<p>We’ve already packed the type into the field number.</p>\n<p>Why stop there?</p>\n<p>What if we could pack the length in too?</p>\n<p>We can.</p>\n<p>now our header can hold</p>\n<p><code>[FIELD NUMBER][LENGTH][TYPE]</code>\nALL in a singular varint!!!! ◝(ᵔᗜᵔ)◜</p>\n<p>Cool. Except for one thing</p>\n<p>its just so much tedious work. To pack a string, I have to tell it ‘this is delimited,’ give it the length, and then give it the bytes. Every. Single. Time.</p>\n<p>PEASANTS DO THAT. Plebeians. we don’t do that.</p>\n<p>So obviously, the solution is to write an entire schema language. We declare our data in a schema file, and let the program deal with all that tedious formatting and encoding nonsense.</p>\n<h2 id=\"building-the-schema-syntax\">Building the schema syntax</h2>\n<p>alright soo… we need to decide on a syntax that doesn’t suck your soul (looking at you protobuf)</p>\n<p>field numbers.. what are they? index right. how do you index stuff in your grocery list? you write number. item why not use the same!</p>\n<p>so something like this:</p>\n<p><code>1. name</code>\nwe need the schema to represent the types. lets just steal how other languages do it and do it like this:</p>\n<p><code>1. name: type</code>\nneat huh. but wait. how do we indicate the end of a message(struct)</p>\n<p>well we could do {} but its not very nice is it. why over complicate stuff its a list. lists have an end. lets have an end.</p>\n<pre><code>message name:\n    1. name:type\n    2. name:type\n    3. name:type\nend</code></pre>\n<p>and no, indentations shouldn’t matter. its so annoying to work with languages where indentation matters. its just painful. lets just not do that.</p>\n<p>so..how will you identify the end of the line without a semicolon? new line char?</p>\n<p>I mean we could do that but what if they wanted to type it in a single line, its ugly but say they want to for whatever reason. lets not restrict that. but semicolons are ugly.</p>\n<p>we could use the number! if we see a number and a dot we just consider it a new entry!</p>\n<p>neat huh.𐔌<em>ˊᵕˋ</em>𐦯</p>\n<p>problem…what if the field is just not found when the compiler is reading it? we could crash…and we would for all the fields but we do want optional fields don’t we. lets just go with the obvious route and do something like this:</p>\n<pre><code>`number. name: type = defaultValue`</code></pre>\n<p>and that’s optional. pretty intuitive.</p>\n<p>now that we are here anyway lets think about all the different kinds of data people could represent…</p>\n<p>well we obviously got our entries of types.</p>\n<p>oh we would need lists. we need some sort of way to pick between a few things so we would need an enum. great lets think about those..</p>\n<p>oh we could just have them be function like, that’s pretty intuitive.</p>\n<pre><code>`number. name: structure(type)`</code></pre>\n<p>like:</p>\n<pre><code>`1. friends: list(People)`</code></pre>\n<p>oh wait but what is people? oh it should be a message too. something like this:</p>\n<pre><code>message People:\n    1. name: string\n    2. location: string\nend</code></pre>\n<p>huh.. we need custom types as well… so in the language do we want to have everything be in order? ehh that’s pretty cringe later we could just do a pass and put together all of our table and just have it refer to that struct.</p>\n<p>hey thanks to this we can have self reference as well. cuz when we pack it it will just be a pointer to the struct! so we can do something like:</p>\n<pre><code>message People:\n    1. name: string\n    2. location: string\n    3. friends: list(People)</code></pre>\n<p>now for enums. cuz we have two passes we can just put enums outside of message! it doesn’t matter where it is declared either! we can declare our enum something like this:</p>\n<pre><code>enum Name:\n    number. name \n    number. name\n    number. name\nend</code></pre>\n<p>also protobuf does this thing where it forces you to have the enum start from 0 like do we REALLY need that? is the compiler so dumb it cant tell? lets just have the compiler handle that by default</p>\n<p>oh wait…. people might need some way to represent an enum but all the types are not the same. yup. that’s a union. lets just add that no problemo</p>\n<pre><code>`number. name: union(type, type, type...)`</code></pre>\n<p>oh wait.. our default value. how would that work with unions. if we have a union like this:</p>\n<pre><code>`number. name: union(i32, i64) = 10`</code></pre>\n<p>it is ambiguous weather 10 is an i32 or an i64 so when field is not found what should we do?</p>\n<p>we could fix this by having the default be next to the type.</p>\n<pre><code>`number. name: union(i32 = 10, i64)`</code></pre>\n<p>now if a the field is not found it will be an i32 with value 10</p>\n<p>okay well people would wanna map stuff as well…so we add</p>\n<pre><code>`number. name: map(key_type, value_type)`</code></pre>\n<p>lets just add syntax for declaring packages and importing stuff:</p>\n<p><code>package &quot;package name&quot;</code><code>import &quot;package&quot;</code>\nhey would you look at that we have some very neat syntax. lets just put it all together:</p>\n<pre><code>package &quot;com.game.core&quot;\nimport &quot;math.jbin&quot;\nenum Activity:\n    1. active\n    2. inactive\nend\nmessage Player:\n    1. name: string\n    2. health: i32 = 100\n    3. weapons: list(string)\n    4. connections: list(Player)\n    5. activeStatus: Activity\n    6. inventory: map(string, i32)\n    7. balance: union(string=&quot;empty&quot;, i32)\nend</code></pre>\n<p>oh wait…we have a whole language now….</p>\n<p>anyway. this is the equivalent of it in .proto</p>\n<pre><code>syntax = &quot;proto2&quot;;\npackage com.game.core;\nimport &quot;math.proto&quot;;\nenum Activity {\n  ACTIVE = 1;\n  INACTIVE = 2;\n}\nmessage Player {\n  optional string name = 1;\n  optional int32 health = 2 [default = 100];\n  \n  repeated string weapons = 3;\n  repeated Player connections = 4;\n  \n  optional Activity active_status = 5;\n  map&lt;string, int32&gt; inventory = 6;\n  \n  oneof balance {\n    string empty = 7;\n    int32 amount = 8;\n  }\n}</code></pre>\n<p>look at that. ew.</p>\n<h2 id=\"time-to-actually-implement-this\">Time to actually implement this</h2>\n<p>To implement this I naturally chose C++.</p>\n<p>By “naturally,” I mean I wanted something I could put on a resume and wasn’t in the mood to fight the Rust borrow checker. I have aged enough.</p>\n<p>Okay. Enough designing. Time to actually make the thing work.</p>\n<p>First, we need to encode our binary.</p>\n<h3 id=\"encoding-and-decoding-our-data\">Encoding and Decoding our Data</h3>\n<p>To encode a varint, it’s basically just this:</p>\n<pre><code>void encodeVariant(std::vector&lt;uint8_t&gt; &amp;buffer, uint64_t value) {\n    while (value &gt;= 128) {\n        uint8_t lower = value &amp; 127; // 127 = 01111111\n        lower |= 128;\n        buffer.push_back(lower);\n        value &gt;&gt;= 7;\n    }\n    uint8_t lower = value &amp; 127;\n    buffer.push_back(lower);\n}</code></pre>\n<p>This builds our varint. If the value is &gt;= 128, we take the lowest 7 bits, set the MSB to 1 to say “there’s more,” and append it to the buffer.</p>\n<p>Then we shift the value by 7 bits and repeat.</p>\n<p>For the final byte, we leave the MSB at 0.</p>\n<p>pretty simple. and we can decode it using this:</p>\n<pre><code>uint64_t decodeVariant(const std::vector&lt;uint8_t&gt; &amp;buffer, size_t &amp;offset) {\n    uint64_t result = 0;\n    int shift = 0;\n    while ((buffer[offset] &amp; 128) != 0 &amp;&amp; offset &lt; buffer.size()) {\n        uint8_t byte = buffer[offset++];\n        byte &amp;= 127;\n        const uint64_t cast = static_cast&lt;uint64_t&gt;(byte) &lt;&lt; shift;\n        result |= cast;\n        shift += 7;\n    }\n    uint8_t byte = buffer[offset++] &amp; 127;\n    const uint64_t cast = static_cast&lt;uint64_t&gt;(byte) &lt;&lt; shift;\n    result |= cast;\n    return result;\n}</code></pre>\n<p>Now that we can encode and decode our varints using LEB128 lets encode and decode some tags(our header metadata we discussed about)</p>\n<pre><code>void encodeTag(std::vector&lt;uint8_t&gt; &amp;buffer, uint32_t fieldNumber,\n               wiretype wiretype) {\n    uint64_t val = (static_cast&lt;uint64_t&gt;(fieldNumber) &lt;&lt; 2);\n    val |= static_cast&lt;uint64_t&gt;(wiretype);\n    encodeVariant(buffer, val);\n}\nvoid decodeTag(const std::vector&lt;uint8_t&gt; &amp;buffer, size_t &amp;offset,\n               uint32_t &amp;outFieldNumber, wiretype &amp;outType) {\n    uint64_t val = decodeVariant(buffer, offset);\n    outType = static_cast&lt;wiretype&gt;(val &amp; 3);\n    outFieldNumber = val &gt;&gt; 2;\n}</code></pre>\n<p>And we can add a couple of helpers for strings:</p>\n<pre><code>void encodeString(std::vector&lt;uint8_t&gt; &amp;buffer, uint32_t fieldNumber,\n                  const std::string &amp;text) {\n    encodeTag(buffer, fieldNumber, wiretype::Delimited);\n    encodeVariant(buffer, text.size());\n    buffer.insert(buffer.end(), text.begin(), text.end());\n}\nstd::string decodeString(const std::vector&lt;uint8_t&gt; &amp;buffer, size_t &amp;offset) {\n    uint64_t size = decodeVariant(buffer, offset);\n    std::string result =\n        std::string(buffer.begin() + offset, buffer.begin() + offset + size);\n    offset += size;\n    return result;\n}</code></pre>\n<p>Encoding a string is now just three things: write its tag, write its length, then write its bytes.</p>\n<p>believe it or not, that’s basically all our core engine done! the rest is just the compiler and the json conversion stuff! ٩(^ᗜ^ )و ´-</p>\n<p>Great! now that we have written our binary encoding… lets make a compiler should be simple…right?</p>\n<h3 id=\"building-compiler\">Building Compiler</h3>\n<p>Yes it is a compiler. stop it I don’t wanna call it a transpiler it is a compiler. the definition is:</p>\n<p>compiler is a computer program that translates source code written in one programming language (the source language) into another programming language (the target language), while preserving the exact meaning and behavior of the original code.</p>\n<p>ours does that. gtfo compiler people. my program takes my schema and translates it into either a binary or code.</p>\n<h4 id=\"the-problem\">The problem</h4>\n<p>Our language is pretty cute. Pretty slick.</p>\n<p>Unfortunately, the computer has no idea what any of it means.</p>\n<p>It’s just bytes.</p>\n<p>so lets assign meaning to these symbols!</p>\n<p>lets take our syntax:</p>\n<pre><code>`1. name:string`</code></pre>\n<p>so its in the structure of:</p>\n<p><code>number → dot → identifier → colon → type</code>\nand then we can use this to build our representation of these entries. that’s what a Lexer does</p>\n<h4 id=\"building-the-lexer\">Building the Lexer</h4>\n<p>Yes lexer is like what you think it is. its just a big loop with a bunch of if statements. All it does is read our file and spit out these tokens.(not the ai kind)</p>\n<pre><code>enum class TokenType {\n    Keyword_Message,\n    Keyword_Enum,\n    Keyword_Optional,\n    Keyword_Map,\n    Keyword_Union,\n    Keyword_End,\n    Keyword_Package,\n    Keyword_Import,\n    Identifier,\n    Number,\n    StringLiteral,\n    Comment,\n    Equals,\n    Colon,\n    Comma,\n    Dot,\n    LParen,\n    RParen,\n    EndOfFile\n};</code></pre>\n<p>so our file</p>\n<pre><code>message User:\n    1. name: string\n    2. id: i32</code></pre>\n<p>the tokenized output would be:</p>\n<pre><code>    [Keyword_Message]\n    [Identifier: &quot;User&quot;]\n    [Colon]\n    [Number: 1]\n    [Dot]\n    [Identifier: &quot;name&quot;]\n    [Colon]\n    [Identifier: &quot;string&quot;]\n    [Number: 2]\n    [Dot]\n    [Identifier: &quot;id&quot;]\n    [Colon]\n    [Identifier: &quot;i32&quot;]\n    [EndOfFile]</code></pre>\n<p>Pretty neat. Now the parser doesn’t have to decipher a bunch of raw characters. It can just work with these tokens. parser looks at these and then actually emits the AST (abstract syntax tree).</p>\n<h4 id=\"what-is-ast\">What is AST?</h4>\n<p>Its just how we represent our program data in graphLang it was a literal tree node that held all the data and all data was just a single shape. For this one we can define our schema as the root. we just have two kinds of messages rn messages and enums so our head can be defined as:</p>\n<pre><code>struct Schema {\n    std::vector&lt;EnumDef&gt; enums;\n    std::vector&lt;MessageDef&gt; messages;\n};</code></pre>\n<p>so our Schema will be the root node and the tree will look like this:</p>\n<pre><code>    (Schema)\n    /      \\\n(enums) (messages)</code></pre>\n<h4 id=\"building-the-ast\">Building the AST</h4>\n<p>soo…. what is enums and messages? from the definition above you can see its a vector of structs. our messages can be defined as:</p>\n<pre><code>struct MessageDef {\n    std::string name;\n    std::vector&lt;Field&gt; fields;\n    int line = 0;\n    std::string comment = &quot;&quot;;\n};</code></pre>\n<p>and then our Field as:</p>\n<pre><code>struct Field {\n    uint32_t number;\n    std::string name;\n    DataType type;\n    bool isOptional = false;\n    std::string defaultValue = &quot;&quot;;\n    int line = 0;\n    std::string comment = &quot;&quot;;\n};</code></pre>\n<p>now our AST looks like this:</p>\n<pre><code>Schema\n└── messages\n    └── MessageDef\n        ├── name\n        ├── line\n        ├── comment\n        └── fields\n            └── Field\n                ├── number\n                ├── name\n                ├── type\n                ├── isOptional\n                ├── defaultValue\n                ├── line\n                └── comment</code></pre>\n<p>and all that’s left is our enum definition:</p>\n<pre><code>struct EnumDef {\n    std::string name;\n    std::vector&lt;EnumEntry&gt; entries;\n    int line = 0;\n    std::string comment = &quot;&quot;;\n};</code></pre>\n<p>and our enum entry:</p>\n<pre><code>struct EnumEntry {\n    uint32_t number;\n    std::string name;\n    int line = 0;\n    std::string comment = &quot;&quot;;\n};</code></pre>\n<p>Great! we got all our pieces. our AST looks like this now:</p>\n<pre><code>Schema\n├── Enums\n│   └── EnumDef\n│       ├── name\n│       └── entries\n│           └── EnumEntry\n│               ├── number\n│               └── name\n│\n└── Messages\n    └── MessageDef\n        ├── name\n        └── fields\n            └── Field\n                ├── number\n                ├── name\n                ├── type\n                ├── isOptional\n                └── defaultValue</code></pre>\n<h4 id=\"building-the-parser\">Building the parser</h4>\n<p>we use this and build our parser. and yes. our parser is just a loop with if statements (well technically recursive decent or whatever but recursion is just a form of iteration).</p>\n<p>it is actually pretty simple the whole loop:</p>\n<pre><code>Schema Parser::parse() {\n    Schema s;\n    while (!isAtEnd()) {\n        if (peek().type == TokenType::Comment) {\n            consume();\n            continue;\n        }\n        if (peek().type == TokenType::Keyword_Package) {\n            consume();\n            if (peek().type == TokenType::StringLiteral) {\n                s.packageName = consume().value;\n            } else {\n                error(&quot;Expected string literal after package&quot;);\n            }\n        } else if (peek().type == TokenType::Keyword_Import) {\n            consume();\n            if (peek().type == TokenType::StringLiteral) {\n                s.imports.push_back(consume().value);\n            } else {\n                error(&quot;Expected string literal after import&quot;);\n            }\n        } else if (peek().type == TokenType::Keyword_Message)\n            s.messages.push_back(parseMessage());\n        else if (peek().type == TokenType::Keyword_Enum)\n            s.enums.push_back(parseEnum());\n        else\n            error(&quot;Unexpected token in global scope&quot;);\n    }\n    return s;\n}</code></pre>\n<p>and for each of the types its just a bunch if conditions checking and then returning the AST node.\n<code>consume()</code> returns the current token and moves the pointer over by one and <code>peek()</code> gives you the next token data without moving the pointer</p>\n<p>GREAT! lets test it out.</p>\n<pre><code>    message User:\n        1. name: adfasdfasdfasdfasd\n        2. id: 012349</code></pre>\n<p>lets see what our parser says:</p>\n<p>“Mighty good mate! seems excellent innit? want a cuppa?”</p>\n<p>yeah..so we need a thing that checks the user isn’t just syntactically correct but also the shit makes sense.</p>\n<p>so we need a fact checker, a twitter community note if you will.</p>\n<p>And that’s what our semantic analyzer does!</p>\n<h4 id=\"building-the-semantic-analyzer\">Building the Semantic Analyzer</h4>\n<p>So to validate stuff we will be doing it in two passes.</p>\n<ul><li>First pass we build all of our symbols and put them into a table. (Basically just all the valid stuff that are identifiable)</li><li>Second pass we use that the table we built validate the AST.</li></ul>\n<pre><code>void SemanticAnalyzer::analyze(const Schema &amp;schema) {\n    buildSymbolTable(schema);\n    validateEnums(schema);\n    validateMessages(schema);\n    validateCyclicDependencies(schema);\n}</code></pre>\n<p>Because we do this in two passes, declaration order doesn’t matter. Which is pretty neat.</p>\n<p>So during <code>validateMessages</code> , when it looks at\n<code>adfasdfasdfasdfasd</code> , it checks our symbol table,\nrealizes that type doesn’t exist, and throws an error . It also checks\nthat you didn’t do something stupid like use field\nnumber  1  twice, or use a list as a map key.</p>\n<p>Great! Look at what we got so far!</p>\n<ul><li>A Lexer that chops everything up into tokens</li><li>A Parser that takes the tokens and builds the AST</li><li>A Semantic Analyzer that validates the AST and throws errors</li><li>An encoder/decoder using LEB128.</li></ul>\n<p>Now all that’s left to do is to do two things.</p>\n<ul><li>Convert JSON files into binary dynamically</li><li>Codegen from the schema so they can use them in their language natively</li></ul>\n<h3 id=\"dynamic-packer\">Dynamic Packer</h3>\n<p>Now for the dynamic packer. It takes a schema, takes a JSON file, and builds our binary.</p>\n<p>For JSON, I just used <code>nlohmann/json</code>. Parsing JSON on top of everything else would’ve been a completely unnecessary side quest.</p>\n<p>So we have our JSON loaded in memory. We have our AST loaded in memory. Now we just walk them together.</p>\n<p>When the JSON parser sees the key “id”: 123 , it doesn’t know what to do. But it asks the AST! The AST says,</p>\n<p>“Oh, id ? That’s field number 2, and it’s an i32 .”</p>\n<p>So the Packer just calls our  <code>encodeTag()</code>  and  <code>encodeVariant()</code>\nfunctions and spits the bytes into our buffer.</p>\n<p>Notice what we didn’t do? We didn’t generate any C++ code. We didn’t compile any wrappers. We just read the JSON, read the schema, and built the binary dynamically.</p>\n<p>take that protobuf</p>\n<p><strong>Lets test it out!</strong></p>\n<pre><code>{\n    &quot;name&quot;: &quot;Pranav&quot;,\n    &quot;health&quot;: 100\n}</code></pre>\n<p>If we save this as JSON, with spaces and quotes, it’s about 35 bytes</p>\n<p>lets pack it in binary.</p>\n<p>and that is…</p>\n<p><strong>NINE BYTES</strong></p>\n<p><strong>hell yeahh look at that!!!</strong></p>\n<p>74.29% REDUCTION</p>\n<p>it would be even higher if we didn’t have strings and stuff. 6 of the 9 bytes are used for the word “Pranav” but besides the point.</p>\n<p>we have reduced our size of storage by a LOT.</p>\n<p>And decoding is way simpler too. The binary already tells us what each field is supposed to be, so the decoder doesn’t have to deal with JSON’s syntax and type representation.</p>\n<p>But what if you don’t want to pay the cost of dynamically looking everything up at runtime? What if you just want normal structs in your language?</p>\n<p>That’s where AOT code generation comes in.</p>\n<h3 id=\"codegen\">Codegen</h3>\n<p>Hmm… how would one generate code? I mean we have our AST so we know what it looks like and what it is semantically so like its just translating that into our language..</p>\n<p>you could write a <code>cCodeGen()</code> func and add it to our class and call.</p>\n<p>if we wanted a generator for python? oh that’s easy! i’ll just add a <code>pythonCodeGen</code>!</p>\n<p>oh I need a <code>jsCodeGen()</code> okay look. too far. why are you writing javascript. but regardless.</p>\n<p>But apparently the customer is always right in their language preferences or whatever. okay now this class is getting a wee bit too bloated for my liking.</p>\n<p>That’s exactly why we need a visitor pattern!</p>\n<h4 id=\"visitor-pattern\">Visitor Pattern</h4>\n<p>So what is a visitor? in hindsight, design pattern cope for ones who’s language doesn not have algebraic datatypes. So why did I use it even though cpp has the <code>std::variant</code>? Idk I read it in an article once. I wanted ot learn about it.</p>\n<p>Anyway. so a visitor pattern is just a lil handshake the caller and callee do. we get type safety from it. in our implementation our caller passes itself in and becomes the visitor. the callee or the acceptor accepts the visitor and does a little func call on the visitor passing itself in triggering a double dispatch. pretty neat..but its just so much mental overhead for a simple problem.</p>\n<p>to accomplish this we will need to edit our structs in the ADT to also hold a function.</p>\n<p><code>void accept(SchemaVisitor &amp;visitor) const;</code>\nand we create an abstract class with a bunch of virtual methods that our generators implement:</p>\n<pre><code>#pragma once\n#include &quot;Schema.hpp&quot;\nclass SchemaVisitor {\n  public:\n    virtual ~SchemaVisitor() = default;\n    virtual void visit(const Schema &amp;s) = 0;\n    virtual void visit(const MessageDef &amp;m) = 0;\n    virtual void visit(const EnumDef &amp;e) = 0;\n    virtual void visit(const EnumEntry &amp;ee) = 0;\n    virtual void visit(const Field &amp;f) = 0;\n};</code></pre>\n<p>so our c generator for example looks like this:</p>\n<pre><code>#pragma once\n#include &quot;SchemaVisitor.hpp&quot;\n#include &lt;ostream&gt;\n#include &lt;string&gt;\nclass CGenerator : public SchemaVisitor {\n  private:\n    std::ostream &amp;out;\n    std::string currentEnumName;\n    const Schema* currentSchema = nullptr;\n  public:\n    CGenerator(std::ostream &amp;outputStream) : out(outputStream) {}\n    void visit(const Schema &amp;schema) override;\n    void visit(const MessageDef &amp;message) override;\n    void visit(const EnumDef &amp;enumDef) override;\n    void visit(const EnumEntry &amp;ee) override;\n    void visit(const Field &amp;field) override;\n};</code></pre>\n<p>so our C generator just overrides these funcs and in these visits it does this:</p>\n<pre><code>void CGenerator::visit(const EnumDef &amp;enumDef) {\n    currentEnumName = enumDef.name;\n    out &lt;&lt; &quot;typedef enum {\\n&quot;;\n    for (const auto &amp;ee : enumDef.entries) {\n        ee.accept(*this);\n    }\n    out &lt;&lt; &quot;} &quot; &lt;&lt; enumDef.name &lt;&lt; &quot;;\\n\\n&quot;;\n}</code></pre>\n<p>when we do ee.accept(*this) it does this:</p>\n<p><code>void EnumEntry::accept(SchemaVisitor &amp;visitor) const { visitor.visit(*this); }</code>\nso it just calls CGenerator.visit(ee) so the visit for environment entries is called in CGenerator.</p>\n<pre><code>void CGenerator::visit(const EnumEntry &amp;ee) {\n    out &lt;&lt; &quot;    &quot; &lt;&lt; currentEnumName &lt;&lt; &quot;_&quot; &lt;&lt; ee.name &lt;&lt; &quot; = &quot; &lt;&lt; ee.number &lt;&lt; &quot;,\\n&quot;;\n}</code></pre>\n<p>and this is done for each type. and that’s how the code is generated.</p>\n<p>GREAT! lets test it out..</p>\n<p>oh wait. we cant. we don’t have a way to do that. we need a cli.</p>\n<h3 id=\"building-cli\">Building CLI</h3>\n<p>For the CLI I used jarro2783/cxxopts cuz I didn’t want to do manual parsing. it will be a fun project both this and the json I will do them sometime else but for now I used these.</p>\n<p>And building the cli was pretty easy from this library you get a bunch of stuff that just works.</p>\n<p>I just wrote my options.</p>\n<pre><code>options.add_options()(&quot;command&quot;, &quot;Command to run (e.g. build, pack)&quot;,\n                              cxxopts::value&lt;std::string&gt;())(\n            &quot;input&quot;, &quot;Input schema file&quot;, cxxopts::value&lt;std::string&gt;())(\n            &quot;o,out&quot;,\n            &quot;Output (target language for build, or output binary file for &quot;\n            &quot;pack)&quot;,\n            cxxopts::value&lt;std::string&gt;())(&quot;j,json&quot;,\n                                           &quot;Input JSON file (for pack command)&quot;,\n                                           cxxopts::value&lt;std::string&gt;())(\n            &quot;m,msg&quot;, &quot;Root message name to pack (for pack command)&quot;,\n            cxxopts::value&lt;std::string&gt;())(&quot;h,help&quot;, &quot;Print usage&quot;);\n        options.parse_positional({&quot;command&quot;, &quot;input&quot;});\n        auto result = options.parse(argc, argv);</code></pre>\n<p>and it works. it was great. and these options were handled in a bunch of if statements(now that I think about it most of this project has just been a loop and a bunch of if statements)</p>\n<pre><code>if (command == &quot;build&quot;) {\n            if (!result.count(&quot;input&quot;)) {\n                std::cerr &lt;&lt; &quot;Error: No input file specified.&quot; &lt;&lt; std::endl;\n                return 1;\n            }\n            if (!result.count(&quot;out&quot;)) {\n                std::cerr &lt;&lt; &quot;Error: --out flag is required (e.g., --out c).&quot;\n                          &lt;&lt; std::endl;\n                return 1;\n            }\n            std::string inputFile = result[&quot;input&quot;].as&lt;std::string&gt;();\n            std::string targetLang = result[&quot;out&quot;].as&lt;std::string&gt;();\n            std::ifstream file(inputFile);\n            if (!file.is_open()) {\n                std::cerr &lt;&lt; &quot;Error: Could not open file &quot; &lt;&lt; inputFile\n                          &lt;&lt; std::endl;\n                return 1;\n            }\n            std::stringstream buffer;\n            buffer &lt;&lt; file.rdbuf();\n            std::string schemaText = buffer.str();\n            std::vector&lt;Token&gt; tokens = tokenize(schemaText);\n            Parser parser(tokens);\n            Schema schema = parser.parse();\n            SemanticAnalyzer analyzer;\n            analyzer.analyze(schema);\n            if (targetLang == &quot;c&quot;) {\n                CGenerator cGen(std::cout);\n                schema.accept(cGen);\n            } else if (targetLang == &quot;py&quot; || targetLang == &quot;python&quot;) {\n                PythonGenerator pyGen(std::cout);\n                schema.accept(pyGen);\n            } else {\n                std::cerr &lt;&lt; &quot;Code generation for &#39;&quot; &lt;&lt; targetLang\n                          &lt;&lt; &quot;&#39; is not supported yet!&quot; &lt;&lt; std::endl;\n            }</code></pre>\n<p>anyhow. lets test it out.</p>\n<p>we will write our schema as this:</p>\n<pre><code>package &quot;com.mmo.game&quot;\nenum Faction:\n    1. alliance\n    2. horde\n    3. neutral\nend\nmessage Vector3:\n    1. x: i32\n    2. y: i32\n    3. z: i32\nend\nmessage InventoryItem:\n    1. itemId: i32\n    2. quantity: i32\n    3. isSoulbound: bool\nend\nmessage Character:\n    1. id: i64\n    2. name: string\n    3. level: i32\n    4. faction: Faction\n    5. position: Vector3\n    6. inventory: list(InventoryItem)\n    7. attributes: map(string, i32)\nend</code></pre>\n<p>and run build with output as c… and..</p>\n<pre><code>#pragma once\n#include &lt;stdint.h&gt;\n#include &lt;stdbool.h&gt;\n#include &lt;string.h&gt;\n#include &lt;stdlib.h&gt;\ntypedef struct Vector3 Vector3;\ntypedef struct InventoryItem InventoryItem;\ntypedef struct Character Character;\ntypedef struct {\n    InventoryItem* data;\n    size_t length;\n    size_t capacity;\n} jbin_list_InventoryItem;\ntypedef struct {\n    char** keys;\n    int32_t* values;\n    size_t length;\n    size_t capacity;\n} jbin_map_char_ptr_int32;\ntypedef enum {\n    Faction_alliance = 1,\n    Faction_horde = 2,\n    Faction_neutral = 3,\n} Faction;\nstruct Vector3 {\n    int32_t x;\n    int32_t y;\n    int32_t z;\n};\nstruct InventoryItem {\n    int32_t itemId;\n    int32_t quantity;\n    bool isSoulbound;\n};\nstruct Character {\n    int64_t id;\n    char* name;\n    int32_t level;\n    Faction faction;\n    Vector3 position;\n    jbin_list_InventoryItem inventory;\n    jbin_map_char_ptr_int32 attributes;\n};</code></pre>\n<p>LOOK AT THAT!! zero dependency left from the program. we provide all the stuff it needs from our schema!!!</p>\n<p>(note: that isn’t the full c file that was generated. there were also a lot of pack/unpack functions for individual messages, setters and getters and encode/decode funcs)</p>\n<p>we can also encode to binary by passing in a json file.</p>\n<pre><code>{\n    &quot;id&quot;: 123456789,\n    &quot;name&quot;: &quot;LeroyJenkins&quot;,\n    &quot;level&quot;: 60,\n    &quot;faction&quot;: &quot;alliance&quot;,\n    &quot;position&quot;: {\n        &quot;x&quot;: 100,\n        &quot;y&quot;: 200,\n        &quot;z&quot;: 300\n    },\n    &quot;inventory&quot;: [\n        { &quot;itemId&quot;: 999, &quot;quantity&quot;: 1, &quot;isSoulbound&quot;: true },\n        { &quot;itemId&quot;: 45, &quot;quantity&quot;: 100, &quot;isSoulbound&quot;: false }\n    ],\n    &quot;attributes&quot;: {\n        &quot;strength&quot;: 120,\n        &quot;agility&quot;: 45\n    }\n}</code></pre>\n<p>and the size of this json is <code>401 bytes</code></p>\n<p>lets pack it into binary. and the file size is..</p>\n<p><code>80 bytes!</code></p>\n<p>EIGHTY PERCENT REDUCTION</p>\n<p><code>80.05%</code></p>\n<h2 id=\"conclusion\">Conclusion</h2>\n<p>So yeah. I started out just wanting to save some JSON out of spite, and I accidentally built a lexer, a parser, an AST, a semantic analyzer, a dynamic binary packer, and a multi-language code generator. And honestly? Packing binary is fun, actually.</p>\n<p>anyhow, checkout the repo.</p>","headings":[{"level":2,"text":"Why would anyone do this?","id":"why-would-anyone-do-this"},{"level":2,"text":"What is binary packing?","id":"what-is-binary-packing"},{"level":2,"text":"Building the schema syntax","id":"building-the-schema-syntax"},{"level":2,"text":"Time to actually implement this","id":"time-to-actually-implement-this"},{"level":3,"text":"Encoding and Decoding our Data","id":"encoding-and-decoding-our-data"},{"level":3,"text":"Building Compiler","id":"building-compiler"},{"level":3,"text":"Dynamic Packer","id":"dynamic-packer"},{"level":3,"text":"Codegen","id":"codegen"},{"level":3,"text":"Building CLI","id":"building-cli"},{"level":2,"text":"Conclusion","id":"conclusion"}]}}