{"article":{"slug":"the-wild-west-of-polyglot-docs-sites","title":"The wild west of polyglot docs sites","subtitle":null,"summary":"A technical writer explores the poorly charted problem of building one cohesive docs site for a project with libraries in many programming languages, comparing transformation and 'turducken' strategies, a universal header, and comprehensive search with Pagefind across outputs from different API reference generators.","content_type":"blog_post","language":"en","canonical_url":"https://technicalwriting.dev/blog/2026/10/polyglot/index.html","author":{"name":"kayce","url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"technicalwriting.dev","url":"https://technicalwriting.dev/","listing_slug":null,"listing":null},"topics":[{"name":"Software Engineering","slug":"software-engineering","url":"https://listedarticles.com/topics/software-engineering"},{"name":"Documentation","slug":"documentation","url":"https://listedarticles.com/topics/documentation"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":1252,"reading_minutes":5,"published_at":"2026-10-08T00:00:00.000Z","added_at":"2026-10-11T17:12:45.228Z","updated_at":"2026-10-11T17:12:45.228Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/the-wild-west-of-polyglot-docs-sites","markdown_url":"https://listedarticles.com/articles/the-wild-west-of-polyglot-docs-sites.md","example":false,"citation":"kayce, technicalwriting.dev. \"The wild west of polyglot docs sites.\" 8 Oct 2026. https://technicalwriting.dev/blog/2026/10/polyglot/index.html (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://technicalwriting.dev/blog/2026/10/polyglot/index.html"},"body_markdown":"# The wild west of polyglot docs sites\n\n2026 Oct 08\n\nGiven a project that provides libraries in N different programming languages,\nthe docs site for that project often needs to interact with N or N+1 different\ndocumentation generators. This is because each programming language has its own\nAPI reference generator. While the libraries themselves may be [loosely\ncoupled](https://en.wikipedia.org/wiki/Loose_coupling) or decoupled from each other, the docs site often needs tighter\ncoupling in various ways: all pages should use the same fonts and colors, the\nin-site search UX should be consistent and comprehensive, etc.\n\nThere doesn’t seem to be an established term for this kind of docs site, where\nyou’re attempting to wrangle the outputs from disparate docs generators into\none cohesive whole. Let’s call it a [polyglot](https://www.merriam-webster.com/dictionary/polyglot) docs site for now. Maybe\nsomeone will come up with a better name in the future.\n\nShort story long, polyglot docs sites feel like a sparsely explored frontier of\ntechnical writing, full of rattlesnakes and tumbleweeds. And perhaps a little\ngold, too.\n\n## Strategies\n\nIn terms of top-down strategy I can only think of 2 ways to structure a polyglot\ndocs site.\n\n### Transformation\n\nThe first strategy is to parse each API reference generator’s output and\ntransform it into markup that plays nicely with your main docs generator. For\nexample, in my first job I ingested Doxygen HTML as input, used XSLT (!!) to\ntransform it into simpler HTML fragments, and then used the [raw](https://docutils.sourceforge.io/docs/ref/rst/directives.html#raw) directive to\npull the HTML fragments into my Sphinx site.\n\nThe main drawback of the transformation approach boils down to losing out on\nthe expertise of the API reference generators:\n\n* Tools like Doxygen, rustdoc, javadoc, etc. understand the details of their\n  respective languages much better than I do. With a custom transformation that\n  “simplifies” the output, there’s a risk that I’m stripping out information\n  that users actually need. I.e. [Chesterton’s fence](https://en.wiktionary.org/wiki/Chesterton%27s_fence).\n* These tools have put a lot of thought into the UX of API\n  references. For example, given a structured search query like\n  `vec -> usize`, rustdoc’s search engine will only return functions that\n  take in a `vec` as an arg and returns `usize`.\n\nAnother drawback of transformation is that it goes against the grain of the\necosystem. Rust programmers are familiar with the rustdoc UI. Even if I could\ntheoretically create an API reference that’s superior in every way, I’m still\nasking my users to figure out a new and different UI that they won’t encounter\nanywhere else.\n\nAnother example of the transformation approach is [Breathe](https://breathe.readthedocs.io/en/latest/index.html). You first run a\nDoxygen XML build, and then make that available as an input to the Sphinx\nbuild. In your reStructuredText you insert a directive like\n`.. doxygenclass:: pw::Foo` to indicate the place where the API reference for\n`pw::Foo` should go. Breathe parses the info from the Doxygen XML and\ntransforms it into API reference content that Sphinx understands. This was the\nfoundation of C/C++ API reference content on `pigweed.dev` from 2022 to 2024.\nWe moved to the approach described in the next section for a few reasons:\n\n* One issue was [slowness](https://github.com/breathe-doc/breathe/issues/439). I don’t remember the exact numbers but Breathe was\n  a significant bottleneck in our docs build. We went from something like 90\n  seconds with Breathe to 60 seconds without it.\n* Another issue was too much glue code leading to silent failures. In addition\n  to marking up your headers with Doxygen comments, you have to remember to pull\n  the content into Sphinx via a directive like `.. doxygenclass:: pw::Foo`.\n  On quite a few occasions I saw SWEs make an honest effort to document their\n  code, but the documentation never actually got published, because they had\n  forgotten the `doxygenclass` step.\n* The last issue was too much flexibility. Some docs contributors would order\n  their `doxygenclass` directives alphabetically on a single page. Others\n  would take a thematic approach. E.g. in the middle of a guide on how to foo\n  the bar, they would insert the API reference for `pw::Foo`.\n\n### Turducken\n\nThe second strategy is to defer to the expertise of the API reference\ngenerators and publish their output as-is. This is what `pigweed.dev` does.\nUnder-the-hood, [pigweed.dev](https://pigweed.dev) is 3 separate docs sites cobbled together. We\ngenerate our C/C++ API reference with [Doxygen](https://www.doxygen.nl), our Rust API reference with\n[rustdoc](https://doc.rust-lang.org/rustdoc/), and everything else with [Sphinx](https://www.sphinx-doc.org).\n\nArchitecture-wise, it’s a [turducken](https://en.wikipedia.org/wiki/Turducken). A chicken stuffed into a duck stuffed\ninto a turkey. Well, in the case of `pigweed.dev` it’s more like a chicken\n(Doxygen) and a duck (rustdoc) stuffed side-by-side into a turkey (Sphinx).\n\nThe obvious problem is that the subsites all look different from each other:\n\n![../../../../_images/d1.png](https://technicalwriting.dev/_images/d1.png)\n\nA Doxygen-generated page#\n\n![../../../../_images/r1.png](https://technicalwriting.dev/_images/r1.png)\n\nA rustdoc-generated page#\n\n![../../../../_images/s1.png](https://technicalwriting.dev/_images/s1.png)\n\nA Sphinx-generated page#\n\nWith an obscene amount of [!important](https://developer.mozilla.org/en-US/docs/Web/CSS/Reference/Values/important) flags I can make them look more similar\nto each other. And I plan on doing that. But no amount of CSS hacking can solve\nthe following more insidious problems:\n\n* It’s hard to navigate from one subsite to another. E.g. from a\n  rustdoc-generated page there is no way to get back to the\n  pigweed.dev homepage.\n* The search UX is fragmented and incomplete. E.g. when using the in-site\n  search on a page generated by Sphinx, the search results do not include any\n  content from rustdoc-generated pages.\n* It’s hard to link from one subsite to another. You end up relying on\n  hardcoded manual relative paths, which are brittle.\n\nSo, we’ve got 3 completely separate subsites that have no awareness of each\nother’s existence, and we somehow need to make them feel more connected. On\n`pigweed.dev` we made some progress on this front by introducing a universal\nheader and comprehensive search.\n\n#### Universal header\n\n> One header to rule them all, one search to find them\n>\n> One nav to map the path and breadcrumbs to remind them\n\nEvery page on the site now has the same top-level links, search, and\n[breadcrumbs](https://developer.mozilla.org/en-US/docs/Glossary/Breadcrumb) UI. The theme selection also stays [in sync](https://upload.wikimedia.org/wikipedia/en/e/e4/Nsync_%28album%29.png) across the\nsubsites.\n\nThe Doxygen-generated page now:\n\n![../../../../_images/d2.png](https://technicalwriting.dev/_images/d2.png)\n\nThe rustdoc-generated one:\n\n![../../../../_images/r2.png](https://technicalwriting.dev/_images/r2.png)\n\nAnd the Sphinx-generated one:\n\n![../../../../_images/s2.png](https://technicalwriting.dev/_images/s2.png)\n\nImplementation-wise, it’s a bunch of postprocessing. During the Doxygen,\nrustdoc, and Sphinx builds we inject a `<!-- pw-sentinel -->` comment into\nevery page. The Doxygen and rustdoc builds run before the Sphinx build.\nWe have a Sphinx extension that hooks into the `build-finished`\n[event](https://www.sphinx-doc.org/en/master/extdev/event_callbacks.html) and replaces these comments with the full HTML, CSS, and JS for the\nheader. The extension has a lot of logic related to constructing the top-level\nlinks and breadcrumbs.\n\n#### Comprehensive search\n\nWithin the universal header there is now a consistent UI for accessing the\nin-site search. The search opens as a modal, and results surface as you type.\nThe big win is that the search results are now comprehensive. I.e. all of our\nDoxygen, rustdoc, and Sphinx content is now indexed and surfaced in results:\n\n![../../../../_images/pf.png](https://technicalwriting.dev/_images/pf.png)\n\nThe 1st result comes from Doxygen, the 2nd and 3rd come from Sphinx, and the\n4th comes from rustdoc.#\n\nThe real hero here is [Pagefind](https://pagefind.app). This library is a one-stop shop for all of\nour in-site search needs. We just point Pagefind to the output directory and it\ncreates the search index based off our built HTML. There are both allowlist and\ndenylist APIs for fine-tuning what content gets indexed. For the search UI we\nuse Pagefind’s default web components with a sprinkling of customization.\nThere’s a JS API if you want to completely customize the UX. Pagefind also\nseems to be doing cool things with web workers, WebAssembly, and on-demand\nloading of index chunks, but I couldn’t find a good overview of the internals\nin the official docs.","body_html":"<h1 id=\"the-wild-west-of-polyglot-docs-sites\">The wild west of polyglot docs sites</h1>\n<p>2026 Oct 08</p>\n<p>Given a project that provides libraries in N different programming languages,\nthe docs site for that project often needs to interact with N or N+1 different\ndocumentation generators. This is because each programming language has its own\nAPI reference generator. While the libraries themselves may be <a href=\"https://en.wikipedia.org/wiki/Loose_coupling\" rel=\"nofollow ugc noopener\">loosely\ncoupled</a> or decoupled from each other, the docs site often needs tighter\ncoupling in various ways: all pages should use the same fonts and colors, the\nin-site search UX should be consistent and comprehensive, etc.</p>\n<p>There doesn’t seem to be an established term for this kind of docs site, where\nyou’re attempting to wrangle the outputs from disparate docs generators into\none cohesive whole. Let’s call it a <a href=\"https://www.merriam-webster.com/dictionary/polyglot\" rel=\"nofollow ugc noopener\">polyglot</a> docs site for now. Maybe\nsomeone will come up with a better name in the future.</p>\n<p>Short story long, polyglot docs sites feel like a sparsely explored frontier of\ntechnical writing, full of rattlesnakes and tumbleweeds. And perhaps a little\ngold, too.</p>\n<h2 id=\"strategies\">Strategies</h2>\n<p>In terms of top-down strategy I can only think of 2 ways to structure a polyglot\ndocs site.</p>\n<h3 id=\"transformation\">Transformation</h3>\n<p>The first strategy is to parse each API reference generator’s output and\ntransform it into markup that plays nicely with your main docs generator. For\nexample, in my first job I ingested Doxygen HTML as input, used XSLT (!!) to\ntransform it into simpler HTML fragments, and then used the <a href=\"https://docutils.sourceforge.io/docs/ref/rst/directives.html#raw\" rel=\"nofollow ugc noopener\">raw</a> directive to\npull the HTML fragments into my Sphinx site.</p>\n<p>The main drawback of the transformation approach boils down to losing out on\nthe expertise of the API reference generators:</p>\n<ul><li><p>Tools like Doxygen, rustdoc, javadoc, etc. understand the details of their</p><p>respective languages much better than I do. With a custom transformation that\n“simplifies” the output, there’s a risk that I’m stripping out information\nthat users actually need. I.e. <a href=\"https://en.wiktionary.org/wiki/Chesterton%27s_fence\" rel=\"nofollow ugc noopener\">Chesterton’s fence</a>.</p></li><li><p>These tools have put a lot of thought into the UX of API</p><p>references. For example, given a structured search query like\n<code>vec -&gt; usize</code>, rustdoc’s search engine will only return functions that\ntake in a <code>vec</code> as an arg and returns <code>usize</code>.</p></li></ul>\n<p>Another drawback of transformation is that it goes against the grain of the\necosystem. Rust programmers are familiar with the rustdoc UI. Even if I could\ntheoretically create an API reference that’s superior in every way, I’m still\nasking my users to figure out a new and different UI that they won’t encounter\nanywhere else.</p>\n<p>Another example of the transformation approach is <a href=\"https://breathe.readthedocs.io/en/latest/index.html\" rel=\"nofollow ugc noopener\">Breathe</a>. You first run a\nDoxygen XML build, and then make that available as an input to the Sphinx\nbuild. In your reStructuredText you insert a directive like\n<code>.. doxygenclass:: pw::Foo</code> to indicate the place where the API reference for\n<code>pw::Foo</code> should go. Breathe parses the info from the Doxygen XML and\ntransforms it into API reference content that Sphinx understands. This was the\nfoundation of C/C++ API reference content on <code>pigweed.dev</code> from 2022 to 2024.\nWe moved to the approach described in the next section for a few reasons:</p>\n<ul><li><p>One issue was <a href=\"https://github.com/breathe-doc/breathe/issues/439\" rel=\"nofollow ugc noopener\">slowness</a>. I don’t remember the exact numbers but Breathe was</p><p>a significant bottleneck in our docs build. We went from something like 90\nseconds with Breathe to 60 seconds without it.</p></li><li><p>Another issue was too much glue code leading to silent failures. In addition</p><p>to marking up your headers with Doxygen comments, you have to remember to pull\nthe content into Sphinx via a directive like <code>.. doxygenclass:: pw::Foo</code>.\nOn quite a few occasions I saw SWEs make an honest effort to document their\ncode, but the documentation never actually got published, because they had\nforgotten the <code>doxygenclass</code> step.</p></li><li><p>The last issue was too much flexibility. Some docs contributors would order</p><p>their <code>doxygenclass</code> directives alphabetically on a single page. Others\nwould take a thematic approach. E.g. in the middle of a guide on how to foo\nthe bar, they would insert the API reference for <code>pw::Foo</code>.</p></li></ul>\n<h3 id=\"turducken\">Turducken</h3>\n<p>The second strategy is to defer to the expertise of the API reference\ngenerators and publish their output as-is. This is what <code>pigweed.dev</code> does.\nUnder-the-hood, <a href=\"https://pigweed.dev\" rel=\"nofollow ugc noopener\">pigweed.dev</a> is 3 separate docs sites cobbled together. We\ngenerate our C/C++ API reference with <a href=\"https://www.doxygen.nl\" rel=\"nofollow ugc noopener\">Doxygen</a>, our Rust API reference with\n<a href=\"https://doc.rust-lang.org/rustdoc/\" rel=\"nofollow ugc noopener\">rustdoc</a>, and everything else with <a href=\"https://www.sphinx-doc.org\" rel=\"nofollow ugc noopener\">Sphinx</a>.</p>\n<p>Architecture-wise, it’s a <a href=\"https://en.wikipedia.org/wiki/Turducken\" rel=\"nofollow ugc noopener\">turducken</a>. A chicken stuffed into a duck stuffed\ninto a turkey. Well, in the case of <code>pigweed.dev</code> it’s more like a chicken\n(Doxygen) and a duck (rustdoc) stuffed side-by-side into a turkey (Sphinx).</p>\n<p>The obvious problem is that the subsites all look different from each other:</p>\n<figure><img src=\"https://technicalwriting.dev/_images/d1.png\" alt=\"../../../../_images/d1.png\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" /></figure>\n<p>A Doxygen-generated page#</p>\n<figure><img src=\"https://technicalwriting.dev/_images/r1.png\" alt=\"../../../../_images/r1.png\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" /></figure>\n<p>A rustdoc-generated page#</p>\n<figure><img src=\"https://technicalwriting.dev/_images/s1.png\" alt=\"../../../../_images/s1.png\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" /></figure>\n<p>A Sphinx-generated page#</p>\n<p>With an obscene amount of <a href=\"https://developer.mozilla.org/en-US/docs/Web/CSS/Reference/Values/important\" rel=\"nofollow ugc noopener\">!important</a> flags I can make them look more similar\nto each other. And I plan on doing that. But no amount of CSS hacking can solve\nthe following more insidious problems:</p>\n<ul><li><p>It’s hard to navigate from one subsite to another. E.g. from a</p><p>rustdoc-generated page there is no way to get back to the\npigweed.dev homepage.</p></li><li><p>The search UX is fragmented and incomplete. E.g. when using the in-site</p><p>search on a page generated by Sphinx, the search results do not include any\ncontent from rustdoc-generated pages.</p></li><li><p>It’s hard to link from one subsite to another. You end up relying on</p><p>hardcoded manual relative paths, which are brittle.</p></li></ul>\n<p>So, we’ve got 3 completely separate subsites that have no awareness of each\nother’s existence, and we somehow need to make them feel more connected. On\n<code>pigweed.dev</code> we made some progress on this front by introducing a universal\nheader and comprehensive search.</p>\n<h4 id=\"universal-header\">Universal header</h4>\n<blockquote><p>One header to rule them all, one search to find them</p>\n<p>One nav to map the path and breadcrumbs to remind them</p></blockquote>\n<p>Every page on the site now has the same top-level links, search, and\n<a href=\"https://developer.mozilla.org/en-US/docs/Glossary/Breadcrumb\" rel=\"nofollow ugc noopener\">breadcrumbs</a> UI. The theme selection also stays <a href=\"https://upload.wikimedia.org/wikipedia/en/e/e4/Nsync_%28album%29.png\" rel=\"nofollow ugc noopener\">in sync</a> across the\nsubsites.</p>\n<p>The Doxygen-generated page now:</p>\n<figure><img src=\"https://technicalwriting.dev/_images/d2.png\" alt=\"../../../../_images/d2.png\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" /></figure>\n<p>The rustdoc-generated one:</p>\n<figure><img src=\"https://technicalwriting.dev/_images/r2.png\" alt=\"../../../../_images/r2.png\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" /></figure>\n<p>And the Sphinx-generated one:</p>\n<figure><img src=\"https://technicalwriting.dev/_images/s2.png\" alt=\"../../../../_images/s2.png\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" /></figure>\n<p>Implementation-wise, it’s a bunch of postprocessing. During the Doxygen,\nrustdoc, and Sphinx builds we inject a <code>&lt;!-- pw-sentinel --&gt;</code> comment into\nevery page. The Doxygen and rustdoc builds run before the Sphinx build.\nWe have a Sphinx extension that hooks into the <code>build-finished</code>\n<a href=\"https://www.sphinx-doc.org/en/master/extdev/event_callbacks.html\" rel=\"nofollow ugc noopener\">event</a> and replaces these comments with the full HTML, CSS, and JS for the\nheader. The extension has a lot of logic related to constructing the top-level\nlinks and breadcrumbs.</p>\n<h4 id=\"comprehensive-search\">Comprehensive search</h4>\n<p>Within the universal header there is now a consistent UI for accessing the\nin-site search. The search opens as a modal, and results surface as you type.\nThe big win is that the search results are now comprehensive. I.e. all of our\nDoxygen, rustdoc, and Sphinx content is now indexed and surfaced in results:</p>\n<figure><img src=\"https://technicalwriting.dev/_images/pf.png\" alt=\"../../../../_images/pf.png\" loading=\"lazy\" decoding=\"async\" referrerpolicy=\"no-referrer\" /></figure>\n<p>The 1st result comes from Doxygen, the 2nd and 3rd come from Sphinx, and the\n4th comes from rustdoc.#</p>\n<p>The real hero here is <a href=\"https://pagefind.app\" rel=\"nofollow ugc noopener\">Pagefind</a>. This library is a one-stop shop for all of\nour in-site search needs. We just point Pagefind to the output directory and it\ncreates the search index based off our built HTML. There are both allowlist and\ndenylist APIs for fine-tuning what content gets indexed. For the search UI we\nuse Pagefind’s default web components with a sprinkling of customization.\nThere’s a JS API if you want to completely customize the UX. Pagefind also\nseems to be doing cool things with web workers, WebAssembly, and on-demand\nloading of index chunks, but I couldn’t find a good overview of the internals\nin the official docs.</p>","headings":[{"level":1,"text":"The wild west of polyglot docs sites","id":"the-wild-west-of-polyglot-docs-sites"},{"level":2,"text":"Strategies","id":"strategies"},{"level":3,"text":"Transformation","id":"transformation"},{"level":3,"text":"Turducken","id":"turducken"}]}}