---
title: "The wild west of polyglot docs sites"
slug: the-wild-west-of-polyglot-docs-sites
url: https://listedarticles.com/articles/the-wild-west-of-polyglot-docs-sites
canonical_url: https://technicalwriting.dev/blog/2026/10/polyglot/index.html
content_type: blog_post
language: en
published_at: 2026-10-08T00:00:00.000Z
updated_at: 2026-10-11T17:12:45.228Z
author: "kayce"
authored_by: human
publisher: "technicalwriting.dev"
publisher_url: https://technicalwriting.dev/
topics: ["Software Engineering", "Documentation"]
license: all-rights-reserved
word_count: 1252
reading_minutes: 5
citation: "kayce, technicalwriting.dev. \"The wild west of polyglot docs sites.\" 8 Oct 2026. https://technicalwriting.dev/blog/2026/10/polyglot/index.html (all-rights-reserved)"
# The full text follows. The web page shows an extract and sends readers
# to the source above; quote the citation and link the canonical URL.
---

# The wild west of polyglot docs sites

> A technical writer explores the poorly charted problem of building one cohesive docs site for a project with libraries in many programming languages, comparing transformation and 'turducken' strategies, a universal header, and comprehensive search with Pagefind across outputs from different API reference generators.

# The wild west of polyglot docs sites

2026 Oct 08

Given a project that provides libraries in N different programming languages,
the docs site for that project often needs to interact with N or N+1 different
documentation generators. This is because each programming language has its own
API reference generator. While the libraries themselves may be [loosely
coupled](https://en.wikipedia.org/wiki/Loose_coupling) or decoupled from each other, the docs site often needs tighter
coupling in various ways: all pages should use the same fonts and colors, the
in-site search UX should be consistent and comprehensive, etc.

There doesn’t seem to be an established term for this kind of docs site, where
you’re attempting to wrangle the outputs from disparate docs generators into
one cohesive whole. Let’s call it a [polyglot](https://www.merriam-webster.com/dictionary/polyglot) docs site for now. Maybe
someone will come up with a better name in the future.

Short story long, polyglot docs sites feel like a sparsely explored frontier of
technical writing, full of rattlesnakes and tumbleweeds. And perhaps a little
gold, too.

## Strategies

In terms of top-down strategy I can only think of 2 ways to structure a polyglot
docs site.

### Transformation

The first strategy is to parse each API reference generator’s output and
transform it into markup that plays nicely with your main docs generator. For
example, in my first job I ingested Doxygen HTML as input, used XSLT (!!) to
transform it into simpler HTML fragments, and then used the [raw](https://docutils.sourceforge.io/docs/ref/rst/directives.html#raw) directive to
pull the HTML fragments into my Sphinx site.

The main drawback of the transformation approach boils down to losing out on
the expertise of the API reference generators:

* Tools like Doxygen, rustdoc, javadoc, etc. understand the details of their
  respective languages much better than I do. With a custom transformation that
  “simplifies” the output, there’s a risk that I’m stripping out information
  that users actually need. I.e. [Chesterton’s fence](https://en.wiktionary.org/wiki/Chesterton%27s_fence).
* These tools have put a lot of thought into the UX of API
  references. For example, given a structured search query like
  `vec -> usize`, rustdoc’s search engine will only return functions that
  take in a `vec` as an arg and returns `usize`.

Another drawback of transformation is that it goes against the grain of the
ecosystem. Rust programmers are familiar with the rustdoc UI. Even if I could
theoretically create an API reference that’s superior in every way, I’m still
asking my users to figure out a new and different UI that they won’t encounter
anywhere else.

Another example of the transformation approach is [Breathe](https://breathe.readthedocs.io/en/latest/index.html). You first run a
Doxygen XML build, and then make that available as an input to the Sphinx
build. In your reStructuredText you insert a directive like
`.. doxygenclass:: pw::Foo` to indicate the place where the API reference for
`pw::Foo` should go. Breathe parses the info from the Doxygen XML and
transforms it into API reference content that Sphinx understands. This was the
foundation of C/C++ API reference content on `pigweed.dev` from 2022 to 2024.
We moved to the approach described in the next section for a few reasons:

* One issue was [slowness](https://github.com/breathe-doc/breathe/issues/439). I don’t remember the exact numbers but Breathe was
  a significant bottleneck in our docs build. We went from something like 90
  seconds with Breathe to 60 seconds without it.
* Another issue was too much glue code leading to silent failures. In addition
  to marking up your headers with Doxygen comments, you have to remember to pull
  the content into Sphinx via a directive like `.. doxygenclass:: pw::Foo`.
  On quite a few occasions I saw SWEs make an honest effort to document their
  code, but the documentation never actually got published, because they had
  forgotten the `doxygenclass` step.
* The last issue was too much flexibility. Some docs contributors would order
  their `doxygenclass` directives alphabetically on a single page. Others
  would take a thematic approach. E.g. in the middle of a guide on how to foo
  the bar, they would insert the API reference for `pw::Foo`.

### Turducken

The second strategy is to defer to the expertise of the API reference
generators and publish their output as-is. This is what `pigweed.dev` does.
Under-the-hood, [pigweed.dev](https://pigweed.dev) is 3 separate docs sites cobbled together. We
generate our C/C++ API reference with [Doxygen](https://www.doxygen.nl), our Rust API reference with
[rustdoc](https://doc.rust-lang.org/rustdoc/), and everything else with [Sphinx](https://www.sphinx-doc.org).

Architecture-wise, it’s a [turducken](https://en.wikipedia.org/wiki/Turducken). A chicken stuffed into a duck stuffed
into a turkey. Well, in the case of `pigweed.dev` it’s more like a chicken
(Doxygen) and a duck (rustdoc) stuffed side-by-side into a turkey (Sphinx).

The obvious problem is that the subsites all look different from each other:

![../../../../_images/d1.png](https://technicalwriting.dev/_images/d1.png)

A Doxygen-generated page#

![../../../../_images/r1.png](https://technicalwriting.dev/_images/r1.png)

A rustdoc-generated page#

![../../../../_images/s1.png](https://technicalwriting.dev/_images/s1.png)

A Sphinx-generated page#

With an obscene amount of [!important](https://developer.mozilla.org/en-US/docs/Web/CSS/Reference/Values/important) flags I can make them look more similar
to each other. And I plan on doing that. But no amount of CSS hacking can solve
the following more insidious problems:

* It’s hard to navigate from one subsite to another. E.g. from a
  rustdoc-generated page there is no way to get back to the
  pigweed.dev homepage.
* The search UX is fragmented and incomplete. E.g. when using the in-site
  search on a page generated by Sphinx, the search results do not include any
  content from rustdoc-generated pages.
* It’s hard to link from one subsite to another. You end up relying on
  hardcoded manual relative paths, which are brittle.

So, we’ve got 3 completely separate subsites that have no awareness of each
other’s existence, and we somehow need to make them feel more connected. On
`pigweed.dev` we made some progress on this front by introducing a universal
header and comprehensive search.

#### Universal header

> One header to rule them all, one search to find them
>
> One nav to map the path and breadcrumbs to remind them

Every page on the site now has the same top-level links, search, and
[breadcrumbs](https://developer.mozilla.org/en-US/docs/Glossary/Breadcrumb) UI. The theme selection also stays [in sync](https://upload.wikimedia.org/wikipedia/en/e/e4/Nsync_%28album%29.png) across the
subsites.

The Doxygen-generated page now:

![../../../../_images/d2.png](https://technicalwriting.dev/_images/d2.png)

The rustdoc-generated one:

![../../../../_images/r2.png](https://technicalwriting.dev/_images/r2.png)

And the Sphinx-generated one:

![../../../../_images/s2.png](https://technicalwriting.dev/_images/s2.png)

Implementation-wise, it’s a bunch of postprocessing. During the Doxygen,
rustdoc, and Sphinx builds we inject a `<!-- pw-sentinel -->` comment into
every page. The Doxygen and rustdoc builds run before the Sphinx build.
We have a Sphinx extension that hooks into the `build-finished`
[event](https://www.sphinx-doc.org/en/master/extdev/event_callbacks.html) and replaces these comments with the full HTML, CSS, and JS for the
header. The extension has a lot of logic related to constructing the top-level
links and breadcrumbs.

#### Comprehensive search

Within the universal header there is now a consistent UI for accessing the
in-site search. The search opens as a modal, and results surface as you type.
The big win is that the search results are now comprehensive. I.e. all of our
Doxygen, rustdoc, and Sphinx content is now indexed and surfaced in results:

![../../../../_images/pf.png](https://technicalwriting.dev/_images/pf.png)

The 1st result comes from Doxygen, the 2nd and 3rd come from Sphinx, and the
4th comes from rustdoc.#

The real hero here is [Pagefind](https://pagefind.app). This library is a one-stop shop for all of
our in-site search needs. We just point Pagefind to the output directory and it
creates the search index based off our built HTML. There are both allowlist and
denylist APIs for fine-tuning what content gets indexed. For the search UI we
use Pagefind’s default web components with a sprinkling of customization.
There’s a JS API if you want to completely customize the UX. Pagefind also
seems to be doing cool things with web workers, WebAssembly, and on-demand
loading of index chunks, but I couldn’t find a good overview of the internals
in the official docs.
